# From Hours to Minutes: The Journey of Fast Finality on Filecoin

- Channel: [ETHBerlin](https://streameth.org/ethberlin)
- Date: 2025-06-19
- Duration: 25:35
- Watch: https://streameth.org/watch/6854224c90bd41297b649d64

## Description

Embark on a behind-the-scenes tour of transitioning Filecoin’s consensus mechanism from taking 7.5 hours to reaching finality, down to just a few minutes, with the adoption of Fast Finality (F3). This talk is all about our experiences and the lessons learned while changing consensus protocols on a live decentralized network. I’ll dive into the real-world challenges we faced, from passively testing new protocols across a large-scale system to gracefully implementing these changes without disrupting the existing setup.

## Transcript

Hi, hello everyone. Thank you so much for being here. Thank you for Tickleberg for having me around again. I really appreciate it. So, I'm, I want to give a talk about an engineer's experience of swapping consensus protocol on a live network. And I, as an engineer, I don't think I get to do this more than once in my career. So, I thought I will give a talk about experiences that we came across because I think probably all networks that last long enough would go through the same process. So, I hope that you get something out of it. So, the Friday that almost broke everything. This is basically the punchline of the talk. But before we get there, I want to talk about how we reached that stage. What lead to that Friday? The stuff that I'm going to talk about, I'm just a presenter here, a very small part of a much, much bigger team that's been working on this for, I think, at least two years. These are just the name of a few people that have been involved in fast finality in Filecoin. Yeah, just to, I want to put it on the slide just to show you that, you know, changing consensus just takes a lot of people, takes a lot of time. What is F3 anyway? So, F3 is fast finality for Filecoin. The QR code will take you to the FIT or the specification of it. And essentially, what fast finality, why fast finality was needed for Filecoin network is that before fast finality, the finality in Filecoin was seven and a half hours. It was a probabilistic model that looked back in basically age of the old file. Epochs in Filecoin are 30 seconds, so that comes to seven and a half hours. And as you can imagine, that comes with a lot of problems for practical users. So, imagine going to the bank and having to wait seven and a half hours to get a confirmation that, you know, your money was transferred. Filecoin itself is a network that provides a decentralized storage and provides proof of space time, it's called, and proof that a copy exists and the So, confirmation is very important in terms of things like data availability, for example, in terms of making, building applications on Filecoin. So, this was a major blocker. Enter F3 and what F3 aims to achieve is to reduce that to about two minutes. So, no pressure. In terms of the scale, the Filecoin network itself has about 2,000 active service providers right now, so that should give you a sort of sense of the scale. At the core of Filecoin F3 itself, there is this consensus protocol we call Gossip BFT. If you're familiar with BFT protocols, you will recognize things like prepare and commit. This is a consensus algorithm that is relying on network spinning. There is a talk again by Alejandro, which I highly recommend having a look that goes deep into this, but I'm going to skim over it very quickly because this talk is more about engineering and things that went wrong and how we fixed it. But to give you a quick overview, you will see a quality phase that exists and you will see a decide phase that exists. So, the way that it works is that there is a quality phase where the entire network start advertising the chain that they want to finalize. And that offers Filecoin very strong censorship properties because everybody has a say in terms of creating a list of all the proposals that could possibly be decided. So, if you are a actor with a lot of power and if your proposal is the only one and then the rest of the network goes to something else, then it basically gives an equal chance for everybody to have their say in terms of the change that's being decided. Then, as soon as the quality finishes, we're going to prepare phase and then prepare phase essentially gets agreement that, okay, this is the chain that everybody is going to vote on. And then we enter the commit phase and then the decide phase. The decide phase is, again, something that is added to the GPBFT, which provides an aggregate of the signatures that are proving that the decision was made. The key advantage of having the decide phase here is that we essentially engineer a side channel for propagating decisions that are made. And that enables things like light clients. So, for example, in order for a node to learn what the finality was, you don't need to participate in F3. You can just listen to a topic that publishes these things we call finality certificates that have a list of aggregate signatures and they can see that the chain or as it's called pipset in Filecoin being finalized and here's the signatures that prove it. There's one extra step there, which is converge. So, I just want to quickly go over it. When we're preparing for something, there might be disagreement. And if there is disagreement in terms of the change that's being decided, then the system enters into converge phase, where we wait for the length of time out of the phases for all the proposals to come in. And then we essentially randomly take the fittest one with a consistent selection algorithm, it's called F3. And we take the proposal, we enter, prepare again, and then we go through the same thing again. And every time we enter converge, essentially the rounds increase. The rounds could increase from converge or commit, this is a small detail, but these rounds can increase perpetually. And by design, F3 should really finalize within one round in vast majority of cases. There's another aspect here, which is the time out between phases. So, similar to BFT, there's time out between phases. And the time out exponentially increases based on the number of rounds. So, what is the challenge? The challenge is very simple. From engineering perspective, we just want to make it look like nothing has happened in the network, except finality has got about 180 times faster. That's the challenge. So, if we make it so that it's boring, then I'll consider it a success from engineering perspective. As you can all imagine, testing a consensus protocol is extremely difficult because you never have a perfect testing environment, and mainnet is always bigger and more complicated than the testing environment. So, how would you go about and test something? In Filecoin, there are two other networks that are spun up and exist, which are called ButterflyNet and CalibrationNet. These two, the ButterflyNet is completely, you know, manually stood up. It's totally permissionful. And CalibrationNet runs about 20 nodes in it. Neither of them have a complexity of the mainnet by far. So, really, testing on those networks just gives you a, you know, vague sense of correctness, and it doesn't test the network, test the protocol fully. What we did was we spun up networks of hundreds of nodes with full stack. We used a Kubernetes cluster to run this. We ran different scenarios, but then again, that thing, first of all, costs, and second of all, you have to really reduce the limits of things like sector size, as it's called in Filecoin, to smaller sizes so that the proofs can be generated very cheaply. You have to fake a lot of interactions in order to get the network up, but that gives you some sort of testing, again, really difficult testing. So, what did we do? We introduced this concept passively. If you have worked in the Web2 world, there's this concept of feature flags. It's very similar to feature flags here, where we essentially wanted to roll out the consensus on the mainnet, but without it affecting anything at all. In terms of rolling it out, so here we have a basic architecture of a Filecoin node. The most popular imposition is called Lotus, by the way, but the details don't matter. What really matters is that there is an F3 module that runs inside the node, and that F3 module would take a manifest or a configuration, that is a list of configurations for the protocol itself, and essentially, it bootstraps the F3 consensus mechanism. It asks for the state on the node to say, give me a proposal to finalize, and then finalizes it, and then tells the node, hey, I finalized it, and that proposal then gets checkpointed inside the chain, and then becomes the basis for the next round. And all of this needs to work with the existing consensus protocol in Filecoin, which is called expected consensus, or EC, so there is a whole bunch of gymnastics and ballet, if you like, inside the implementation of this to make sure these two play nice, because really, finality could change in two different ways. So as an engineering team, we had permission access in order to swap the manifest configuration inside the node. We essentially run the whole protocol in full scale, or in gradual increase of the scales across the entire mainnet, and change the manifest, and we engineered a whole bunch of layers to make sure that we didn't actually destroy any existing state, or destroy the F3 module completely when we are reinitializing a new test network. So a lot of work went into getting that correct, but it was an extremely effective way of testing something, because we could use the real chain, we could use the real behavior, we could stand up a test network and leave it running for a week, and see what happens. What it shows us was extremely effective, so I recommend that approach. Fast forward six months. So we have implemented the passive testing, we have implemented the core protocol, we have integrated it, everything is ready to go. Passive testing looked fantastic on calibration net, the numbers look great, we were very happy with the performance of the network, there is a parameter called delta in GPBFD, which is essentially the delay in roundtrip, sorry, it's a roundtrip latency, and there's a massive propagation in the network, and this is a parameter that needs to be measured. So we used the previous value that was used by the current implementation, and everything looked great, fantastic. And then we went ahead and activated S3 on calibration net, because everything looked great, and then we moved on to testing it and something interesting. We realized that the level of bandwidth that the nodes are using, F3 implementation is using, is just well above the layer that is tolerated by the network. So we knew that the bandwidth is going to be increased, because of sheer amount of traffic that's going to be generated by the nodes, so just to give you a point of reference, before F3, there was only a few messages per second being exchanged between the nodes, the transport mechanism that's used by Filecoin is Gossip Sub, which is a DP2P protocol, same protocol that's used by Ethereum, but F3 was really pushing this protocol to its limits. After F3, we're talking about about 250 messages per second more on top of the existing mechanism that was used for only propagating a few, and you can imagine how that can affect the message propagation time across the network. What we observed was that even during slow ramp-up of this passive testing, when we reached about 80% of the network scale, we gradually increased it, we saw a huge spike in bandwidth consumption. In the beginning, there's a bootstrap mechanism in F3, and then there's a steady state. So in the bootstrap mechanism, the bandwidth usage for nodes increased up to 10x, and then after that, in the steady state, even at 80% scale, we were talking about 4 megabytes per second bandwidth forever, for operators that are used to few, maybe tens of gigabytes of bandwidth usage per second. And that was a big pill to swallow for providers, so we learned. There was another side effect of this thing. It wasn't just a social work of kind of telling people that, hey, this is not a bargain, this is a new thing, but it gives you fast finality, you've got to get used to this, you've got to spend money to get something back, essentially. There was also another effect, which was, because the message propagation was slow, it was causing asynchrony in the network. And as a result of it, if you remember, I mentioned BFD or Gusset BFD, they rely on synchronous communications. So ideally, all the nodes should be in the same phase at all times, and there's timeouts and delays or dynamic adjustments of this latency to at least optimistically try and get nodes to be at the same level. What was happening was that because of the increased bandwidth usage, this synchrony was knocked off. So some nodes were slightly ahead, some nodes were slightly behind, and by the time where a timeout was occurring for some nodes, the other ones would just reach the timeout, and then they would realize, oh, nobody agreed or I didn't receive any messages, and then they will try and decide on base, which is decide on essentially nothing. So that then increased the number of rounds that would require to finalize F3, which then increased the time to finality, which defeats the whole purpose of having fast finality. So how did we solve this? We looked at the messages, and the biggest thing in the messages is the chain that's being finalized. In Filecoin, you have this concept of tipsets, and each tipset has a set of tipset keys, essentially. Each tipset has a tipset key, which is made up of many, many CIDs or content IDs or hashes, essentially. And this thing could be very sizable. This chain was included in every GPBSD message. So in every phase, every update, everybody tells everybody else the chain that they're trying to send out. So what we did was we essentially took that chain out and made it into a side channel lookup. So rather than telling everybody the chain that we're trying to finalize, we tell them a cryptographic key of it. And that way, the message would be much, much smaller, but then you have another problem of telling everybody what key corresponds to what chain, which is a separate thing I'll talk about in a minute. The other thing was it was surprisingly fruitful to use compression, by surprise. So we used ZSD compression. It is an extremely fast algorithm. It is extremely cheap. It's just a fantastic algorithm. And that reduced the size of messages. The compression rate was about 1.7 on average. So the way that it works is that we have a chain exchange. We have a separate topic for exchanging the chain at the top. So what happens is that a node will come up. So at the bottom, I have a timeline. So as you move right, the time progresses. As soon as the node comes up, it finds this proposal and publishes it on this chain and then periodically keeps publishing it on the chain. And it also uses the same chain to discover other people's proposals. Then what happens is that it usually just starts GPBFC phases as normal. And as it publishes a message, we call them now partial messages because they only contain keys. When it receives a partial message, it just buffers it. And there is a best effort, a clever coding to make sure we understand, we can infer messages just using local state. But if we can't infer messages just using a local state, then we rely on chain exchange to tell us what the key corresponds to what chain. Then eventually, it realizes that a key X maps the value Y. And suddenly, all those partial messages become un-partial anymore, become un-partial and become consumable. And then they all get passed down into GPBFC, which is the last bit in the image that you can see. And GPBFC progresses, and that's great. And that reduced the size of messages by 70% and compression rate really helped. So fantastic. Happy faces. Another problem was if anybody has done network operating, they know that it's a really painful thing to do in the permissionless network because you essentially have to coordinate a lot of social aspects in order to roll out a new software update. And everybody has to roughly upgrade in the same time. In S3, there's a set of parameters that need to be measured empirically. And then once those are measured, then S3 can be activated. What we really wanted to avoid is to have this two-step upgrade because you have to measure it, but then you have to release another software and do a network upgrade to actually activate it. And that was a real pain. It's very slow and time-consuming. The trick that we used there was using this self-locking one-time contract manifest. Essentially, we made a release of a software that had a built-in contract address that essentially defined a way by which you could publish a manifest and have that locked in place. And once the manifest is published, then the software itself disables that dynamic manifest that I mentioned earlier. Here's a quick diagram of it. So we have, if you remember, the permissioned engineering team access to dynamic manifest. As soon as a set of committees, which were made up of implementers of most popular FileFront clients, receive a proposal from us to say, okay, this is the parameters that we have tried. You can see the measurements that we got from passive testing. If it looks good to you, please sign this contract. Then they go and sign the contract, and then the code in each of the nodes pulls the state of that contract, say, oh, I've got a new manifest. I'm going to go into lock mode. I've got the final manifest. Then there's a locking period of 96 hours for the community to reflect that and see the changes. And then that manifest would then have the final activation epoch for the entire network. And the network activates, and it's beautiful because we have a single network upgrade, and we've done our passive testing, and we've done our activation in one step, which is great. And then we entered an intense period of testing because we wanted to make sure that we stick to this one network upgrade. It was paramount to make sure there's not that many bugs and so on. So a lot of simulation on, we use Kubernetes cluster. We use modern tools like ChaosMesh. I'm not sure if you've used them. They can fake things like IO failures and network partitions and so on. And we managed to find and fix a lot of bugs as a result of that. So it was a great investment of time. I highly recommend doing that sort of thing. It can be very fruitful. April 29, 2025. Activation day. Great. We're all excited. The contract is locked in place. There's nothing else we can really do because we're locked out. Fantastic. A weight off our shoulder. We've been anxious about it. The time comes and everything goes beautifully boring. The network upgrades, F3 kicks in, plays really nicely with EC. Fantastic. And then we were in the state of disbelief because we had a lot of problems implementing this and had a lot of challenges. We couldn't believe that it's finished. We just couldn't allow ourselves to rest. And then it came the Friday of that week, which was about three days after activation, at instance 6017, which will forever be hacked in my memory. It was Friday evening my time at about six o'clock. I was finally believing that F3 is over, closed my laptop, went to have a drink, and then realized on Saturday that the network is completely stopped. It was total radio silence. Nothing was happening in the network. And it just wasn't clear why. It was as if F3 is not running at all. Yeah. And by the time we got to it, we got to round 11. And remember that exponentially increasing timeout? That meant that it will be hours before we see any sort of interaction in F3 module to even understand what is going on. And that was extremely challenging. So why did that happen? We spent Saturday and Sunday trying to do detective work and realized why. So GossipBFT protocol has this mechanism for rebroadcast. It rebroadcasts messages because the message delivery, it's optimistic. I think in previous talk, it alluded to different types of broadcast. GossipBFT relies on optimistic message delivery. There is no guaranteed message delivery in GossipSub. So we rebroadcast messages to make sure that at least the state we're trying to decide is propagated in the network. The rules of rebroadcast in GPBFT are simple. If there is no progress, then rebroadcast. And lack of progress is defined as there is no increment in instance ID, which is monotonically increasing. There is no change in phases. And there is no change in round. Excellent. That seems like F3 is not progressing, so let's rebroadcast. What was happening was that the rounds were progressing, but not the rest of the state. So the rounds kept progressing. We went to round one, round two, round three, and we entered things like, we entered a state of lack of synchrony where everybody was voting for nothing, which is base. That's comets for bottom. That's what it means. And then we were entering converge. Converge is the stage that it always waits for the total timeout of the state. And then we go to prepare lack of synchrony again. We go to comet to bottom. Another converge. Round increases, and timeout increases. And you can imagine this exponentially increasing timeout, which started from two seconds. And by the time we got to it, it was an order of hours. And rebroadcast never kicks in because the rounds increment. Because as is defined by the implementation, round increasing is progress, right? It technically is progress, but it's not useful progress. So that's a problem that we had. Just to show you what the network looks like, you can see on the far left the amounts of bandwidth, which, by the way, with the bandwidth optimizations that we did, we managed to do one megabyte per second, getting finality, which is fantastic. And then, basically, the throughput, the bandwidth usage completely dropped. We reached 10,000 epochs. We stopped 10,000 epochs behind head, which was fantastically slow. We figured out a problem. We implemented this thing we called rebroadcast aider, which essentially consumes all the messages and then rebroadcasts them. And we managed to get the network back up again. Yay, back to beautiful boring. Today is a month that F3 has been active on maintenance and is performing really well. So this shows you 95th and 99th percent distance from head. So we are about five epochs behind on the 95th percent, which is fantastic. So what do I tell myself if I want to do this again? Bandwidth is a silent killer. It's a reinforcing feedback loop, so it can have cascading effects, so be very careful about bandwidth. Progress is not always progress, so make sure you have a very clear definition of progress. One upgrade is always better than two, so keep that in mind. And compression could be something that is a free performance game, almost, which is magical. Just try compression. It's really easy to write benchmarks for it. Well, the real victory for us was the Filecoin network is now 180x faster in terms of finality at one megabyte per second. Bandwidth uses, which is fantastic, great achievement to the team, and we have learned a lot from it. And I hope that you learned something from our experience. And yeah, thank you for listening. Thank you very much, Mazi. I think we're running a little bit late, so I don't think we can get to the questions. Or let me ask you one question, because I see this one got two upvotes, and then we'll do a short break and go to the next one. One quick question, is there client diversity for Filecoin on the consensus layer? Yes, so there are right now three different client implementations that are using the same consensus protocol. Perfect, thank you. I see there are three more questions. We can't get to them, since we need a little bit of a break in between sessions. So, if I didn't get to your question, feel free to approach Mazi, and he can answer it in person.
