State Archival: Expiring the State Without Telling Anyone - Guillaume Ballet | geth
ETH Belgrade Community·Tue, Oct 6, 2026, 12:00 AM
Speaker
Transcript
Uh, is it working? I Yeah. Okay. Uh, right. So, um, I'm going to talk about a little experiment I run earlier this year.
And, um, yeah, it's, um, it's not the most original work I've done. Um there's some aspects of it that have been done by other clients like uh like Arrigan like uh like uh Nethermine but um it gives uh a view inside the the work of a core developer and uh it also uh uh creates uh like gives rise to a few questions that I think are are interesting for for the future of of the way Ethereum um upgrades. Um so to bring everybody on the same page, uh I just want to explain a bit how the the data that Ethereum holds like all the nons, the contracts, the the ETH, uh they're stored in a big tree and that tree is stored in a database and unfortunately the database is not just like writing uh you know text into a text file. It grows bigger. as it grows bigger, you need to update some index indexes like there's a lot of site structures that will slow uh that needs to be updated and as a result the the performance degrades a bit.
Um and it's a bit of a shame because if you look at how much uh of the state like what share of the state gets gets touched over the course of one month, it's barely 3%. Uh some of it is uh uh new new state that gets added. The the state keeps growing and some of it is just uh old stuff that that gets updated. For example, you spend uh you spend some ETH. Uh that's that's part of it.
However, most of it is completely untouched and you know, some of it is because uh it's people are are huddling. Um, so they're they're not going to touch their their ease for like a year, maybe two uh maybe 10. But uh but some of it is also uh work that has been uh done by me searchers just to try to sandwich transactions. So it's just contracts that get created and will never get touched. So a lot of it can actually be uh taken out of the database.
And there's been uh already some similar work that has been done for uh for history expiry. Uh so history expiry means deleting the old blocks. You have the blockchain it you keep adding block and you want to put those blocks in the database around the head like you can have a reorg. So potentially especially in proof of stake you have uh potentially multiple branches. So you can't just store everything linearly you uh in in a flat file.
you need to be able to say, "Oh, actually this uh I thought this was going to be the head, but actually this branch was going to be the head." So, you can't really store things uh inside a inside a database. Uh sorry, you do need a database to store these things. But after a while, especially in proof of stake where things get uh linear uh sorry finalized, you know that you're never going to be uh going back all the way up to Genesis. So you can afford to move this thing out of the database and uh and put it in what what is called an archive.
So uh just a flat file. And if you need for whatever reason there's a catastrophic bug, you need to uh you need to unfurl the entire database. It's okay because at this point the chain is linear. So you can read start reading from the the end of a of a file and go back to the to the start. Um and with this we freed about 400 gigabytes of the of the database.
Um and the performance uh was noticeable. Uh so let's do the same thing with with state. Um the the Ethereum state. So all the once again all the all the uh all the ERC20s all the all the accounts are stored inside a tree. It's a Merkel tree.
So uh you have the actual data at the bottom of the of the tree and then you the parent is hashing those leaves and then the parent is hashing the parent and so on all the way up. Uh so you actually have a two-dimensional optimization uh uh structure. Uh but everything depends on the the bottom layer. So you can actually delete a lot of uh structure vert sorry a lot of data going vertical. Um, so what does that look like?
Uh, you see on the left you have a tree and the the the size of the tree, I'm going to explain why, but the maximum amount of uh nodes we're going to to delete uh is is a height of three. And the um and you know, if you delete all the internal nodes inside this sub tree and you keep uh you move the the data to a the actual leaf data to an archive, uh you can always reconstruct it. Uh so so those trees are not needed. Why do we limit it to three? Because as soon as you need to resurrect um you will have to rehash everything all that data.
Yes, it can be regenerated but it takes a bit of time and a depth of three is what is acceptable in terms of computation. It's not that much. It can be done quickly. Uh so um so yeah that's that's why three has been chosen. If I hope it's readable.
Um if you need to read that's fine. You don't even need to reconstruct the tree, right? It's uh uh you just go to the to the archive. You you go get your value and and that's it. But if you need to uh to reconstruct uh if you need to update, that's when you need to cryptograph cryptographically commit to everything.
So you will need to rebuild every layer and reinsert them to into the database. And the biggest problem is not so much the hashing, it's the actual reinsertion inside the the database. Um, we could of course try to only reinsert what we actually need instead of reinserting the entire subree, but uh that was an experiment. So, so I haven't done it so far. Um, yeah, TLDDR, I'm going to gloss over that because uh it's it's a recap.
Storing data in a file is a lot uh it's cheaper, it's faster. storing it in the database has brings you some features uh but uh you you have to pay especially for the for the IO for writing to disk and things like this. So let's go to the actual experiment. Uh I initially started with uh with uh hoodie I did a full sync. I stopped halfway through the history of the chain.
I expired all the nodes and then I replayed uh the rest of the chain. And what I could see is that on this uh second uh leg I was a lot faster. I was uh I was like the the gas per second uh was almost uh 1/5 faster. So the performance was was great. What was uh the the extra cost in terms of dispace that I had to pay?
That was roughly uh I forgot the exact detail but it was two 2%. So you would say why uh you know you delete data from the database why is it actually taking more space and the reason for this is because um this is not a very optimized like the archive is not very optimized. The database compresses the data the archive does not. So just uh just that you you lose some uh you lose some uh some space and in the database you also need to store some references. So um so because of that uh you're actually not you're actually taking a bit more space but um yeah I would say for a for a first uh for a least effort approach that's that's pretty good.
Uh then I uh created I did the same thing on mainnet and uh the the picture is quite different. So I was able to delete about a quarter of the data. So the the database size shrunk by by one quarter. Uh but of course I have a 98 GB archive on the on the side. So that means I all in all I save like 3 GB.
Um and that sounds like uh not much. So it's a bit underwhelming especially when you look at the performance it's a bit faster but it's not that much faster. And uh there are two reasons for this. Uh the first one is uh if you look at hoodie the state of hoodie remains very very uh static like uh hoodi doesn't actually write to three uh I mean it keeps writing to the same locations basically whereas mainnet the the 3% over time they move and they cover a lot more a lot more uh space like like the coverage inside the the state is is a lot bigger. So that means you you need to resurrect a lot more.
So part of the problem is the resurrection. The other problem and that's the interesting part is that uh because the database itself uh uses a tree. So you you store a tree inside a tree like uh I call that a tree seption and you have u like the because it's a tree the the storage size uh like based on how much data you store the depth of your tree is kind of a logarithism. I mean it totally is a logarithism and the logarithm starts pretty steep and then it tapers. It it flattens and so that means that if you look uh yeah like the picture is a bit hard to see from from the room but if you look at uh the black spots like uh this is the this represents the uh the the sizes when you're not at um when you you didn't archive.
Uh so the the left most is a hoodie, the rightmost is uh is mainet. And by the way, it's not it's not to scale. Um but if you look at the green dots, they represent the archive. So when you start deleting data from the database and you can see that the hoodie one actually benefits from the from the from the steepness of the curve, but the the archive one not so much. So this is the second reason why you don't really get performance boost yet because the whole uh the whole interesting part about this is not what happens today.
It's long-term uh long-term uh prospects. We today the state is about 300 gigabytes. Uh we want to bring it to 5 terabytes. This is going to mean that even though the logarithm curves increases slowly, it does increase. it goes to infinity eventually.
So what you want to do is protect against the future degradation that's going to happen and this approach would solve this as evidenced by hoody but um yeah today the effect is is minimal. Um once uh all that is said and done uh so far what I've done is just look uh you know if I'm running this experiment nobody knows I'm I'm doing it like uh in terms of consensus no one has has any idea was uh running this on my client. Um we could uh add an extra EIP to actually store the age of the last time a location was touched in the in the state. So that means every time a block is uh sorry an account is uh is written to you would add uh the information for the for when it was written and this way it can be a bit more optimal uh to yeah to expire things because my method was to expire everything and then um see what gets resurrected and that's okay that's uh that that provides interesting results. It does provide a minimal speed up but I could have done uh a lot better by just writing down what was uh written recently keep and keeping that in the database and just expiring what hasn't been touched in uh in over a year and if I keep this process going I would keep the database small and I would uh I would get um I would get that uh that performance like to to endure over time and not having to archive on the site every every time.
Um, okay, I'm going to skip this one because I think I'm running out of time. Um, but yeah, there's been this uh like if you want to go through like we can push this a bit further and uh go with um state expiry. uh Ethereum has tried to achieve state expiry like actually losing state like deleting state from clients for uh you know as far as I remember that was 2018 it was probably the case before um and uh what you would do is just you have the reference to to the archive you keep the reference to the archive but you delete the archive this is what we did for the the blockchain like history expiry we have the reference to the previous blocks um through the cryptography commitment through the hash of the parent block. But we deleted this. Most clients deleted this and of course we have to uh recover that data somehow.
Make sure it's recoverable. That's what Milos for example was working on. Um but uh we could we could build a similar system like it's a bit unwavy but when I think about it uh it's it seems doable. So we could achieve uh history expiry in a very uh in a very uh subtle manner. Initially not doing any change to the protocol which is always the biggest slowdown and um and then uh introduce little by little a very tiny change in the protocol that will be tested by the time we do it.
So everybody will agree there will not be any bickering any disagreement. Everybody will know what needs to be done and uh and when we do this uh we can have history expiry shipped even uh without having a single debate and that's uh my uh my last slide my my conclusion uh you know we could we could take maybe uh this approach and apply it to everything. So I don't know if uh you guys have seen this this picture it's it's called the straw map. It was proposed by uh Ethereum Research and uh it's a five-year plan. So I don't know what you think of five-year plans, but uh none of them ever uh ever come to f to fruition as far as I know.
Um and uh it's full of EIPs and some of them are EIPs that we ship because we think that in five years we will need it except the plan in five years uh keep moving. researchers can't help researching. So they come up with new ideas and that means whatever you do today is for something that might not like is for a moving target. So it's a bit uh it's a bit nwavy. There's a lot of debates.
We spend a lot of time bickering on a core dev which is the call that we have for discussing EIPs. Um and uh and I that's not uh that's not the the right um the right way to proceed. I think everybody is getting a bit burnt out for from this process. Um this is by the way produced by researchers. We get very little user input.
No one, you know, you have ZK, you might have heard of ZK, you might have heard of uh binary trees. You might have heard of formal verification. The user doesn't care about this. Um what the user wants is something that goes fast that when they send a transaction, it's safe and it's uh and it's fast. The rest it's a detail.
Um so I think um we should get more people on ACD that come and say this is what we want and instead of going from a research top down approach we go to a bottoms up appro approach. So I've seen a lot of people advocate for um for uh dictator like benevolent dictators a bit like the Linux kernel does. Uh but I think you know we have this decentralized ethos which we should try to keep it. We just need to do it right with the right kind of discipline. And if we uh actually do our experiments based on what um on what users tell us and we experiment out of protocol first and only when everybody has implemented inside the like the right thing inside their protocol.
We try to enshrine what I think is useful for everybody. Then we have a much smoother uh forking process. uh we can have a lot a lot of faster forks and maybe we don't and that's fine because most of the time is spent working on the clients instead of implementing EIPs that will never pan out. Um so food for thought if uh if I convinced you please support your local core developer by uh telling so on Twitter or coming to ACD and uh and making your voice heard. And with that uh that's the end of my talk.
So thank you very much. And if you have questions, I'm happy to answer them.
Automatic transcript — names and jargon may be misspelled.