Stateless Ethereum Clients - Milos Stankovic | Ethereum Foundation
ETH Belgrade Community·Tue, Oct 7, 2025, 12:00 AM
Stateless Ethereum Clients - Milos Stankovic | Ethereum Foundation
Transcript
Good morning. Um, thank you for the nice introduction. My name is Milos. I work for Ethereum Foundation. And as it was said, I will give a a sneak peek of what the Ethereum Robbuck is, what we are working on and where this is all leading or at least one of the directions.
Um okay so just to know a bit more about my audience um who here never heard about statelessness or who heard about statelessness but doesn't really know what it is refers to what it is about or anything like that. So please raise your hand if you don't know what this talk is actually going to be about. Okay some hands. And who here doesn't know what are Ethereum clients or never heard about Ethereum clients or doesn't know what what even that is? Okay, so everybody at least somewhat familiar with Ethereum clients.
But let's look a bit into more details about it. Anyway, so currently if you discuss about Ethereum clients, they are divided mostly in three categories. Um archive nodes, full nodes and light clients. And I will go into details about them. But let's see first what the Ethereum clients can do.
And I'm mostly talking in the context of execution layer. There's also consensus layer and clients there. But for this talk we mostly talk about execution layer clients. So let's say what they can do that is the column that is the rows and we have the main category that every that every blockchain needs to do is blick menu. And we will see not all clients can do that.
Uh another thing is we want nodes to or clients to verify the blocks that are being built. We want uh to be able to create transactions even if they're not building blocks to maybe create transaction that will be sent to other blocks to other clients to to build them. Um head state refers to accessing uh the the state of the blockchain at the head of the chain. Meaning if you want to know what is my current balance right now at the very head of the chain or either balance or what tokens I own etc. what is the price of the ether asset on the unis swap exchange etc.
So that all refers at the moment and the historical state refers to what was my balance a year ago or something like that and that is historical state refers to very crude access to that state. It doesn't even refers to like for example if you want to answer question give me all transactions from my wallet like that requires some processing and indexing data to be effectively accessible like if you have all historical state you can find that information but it's usually not efficient so you need additional pro prep pre-processing of the data and I'm not even covering that because none of these clients actually support that kind of indexing natively you need to build on top of it uh and requirements is basically a bas comparison between these clients and what they are offering. So the archive nodes they can do all of this but their requirements is depending on the implementation because there are for each of these nodes there are multiple implementations written in different programming languages by different teams. Um so it kind of like depends but usually the requirements if if you want everything you need around 10 terabytes of data which is quite a lot if you if you just want to run it on your own machine. um it's still manageable but kind of quite a heavy heavy requirements to to run that.
Full nodes are the more common type of nodes that people are running. Uh they offer pretty much the same except access to historical data. Um it's maybe ironically they are called full nodes but not offering everything. Might be counterintuitive. The fact is that they actually store all the historical blocks.
So they can create the historical state and give access to that like because they have all the blocks but they don't have the state historical state they could technically recreate the state and give you what was your balance a year ago but it would be very costly for them to do it. So they are full nodes because they have everything that they can do everything but some of the stuff primarily the historical state would be very expensive so they just decide not to do it because it's too expensive operation and you should do something else if you need that. Um light clients well there is no clear definition or requirements of what light client exactly is. There is some kind of consensus that when people talk about flight clients they talk about some kind of software that is not always running or always sing to the head doesn't require a lot of storage like full nodes they require less than two terabytes that's usually uses a benchmark of how much they need light clients expected to require significantly less maybe in order of hundreds of megabytes or something like that but there is no strong requirement of how they operate what exactly functionality they provide etc and there are a couple of attempts of making them more user friendly etc. But it's quite difficult and that's why we have something like this.
They definitely cannot build blocks or verify blocks because in order to do those at the moment you need the entire state and entire state is already uh 400 GB of data or something like that depends on the implementation and light plants by default they are supposed to store significantly less. So they cannot do the first two things. The others depends on how they're implemented. They can do some of it with certain trust assumptions because even for the others you need some kind of access to the state and like clients maybe have them or maybe they use some centralized entity to get them. Maybe they verify them, maybe don't verify them.
So it's kind of questionable what they can do and how much trust assumptions you put there into doing the operations. And what is the goal? The goal is to make the light clients uh much better and much more efficient. basically not being able to to need all the state to do certain functionalities as Ethereum client and that's where the statelessness comes in a road map of Ethereum where we are aiming at. So how would how would that look like?
So first main thing that we want to do is we want block verifiers that they don't need the entire state in order to verify the block. That's the first requirement. So we want everybody to be able to easily run a note and verify the blockchain without actually having the entire state. Second aspect is we probably still need block builders to need to have the entire state. This is a concept called weak statelessness.
There is also another concept called strong statelessness where even the block builders don't need to have the state. um it's a bit more challenging obviously to do that and requires um requires more work. There's not so much research being done in this direction. So I will probably be speaking in the direction of the weak statelessness because that's where the most work and most research can been done so far. Um okay so what are the benefits of the statelessness?
Well, if you have uh if you can run the validators on much smaller machines, um you have much better experience for validators. You can run it on um either a laptop or maybe even a phone or just like any any random device like it probably also doesn't have to be uh all the time online. If you're running a staking validator, then yes, it has to be. But if you just want to uh run a validator to just verify the blockchain, then you don't need to be online all the time, etc. You don't need to sync the state because you don't need the state to begin with.
Um that also helps with scalability because it's much easier to run these kind of clients with then they are easier to run then it's more decentralized meaning everybody can run them. Um it allows a lot of new features that you can do on top of it because it becomes to enable statelessness. It allows making the proofs about the state much easier. That's what requirement we will come to that. It also helps with layer one and layer two interoperability because it becomes easier to prove the state and all the transition etc.
So what is the main idea behind the statelessness? So if we have blocks, every block has certain fields, hash of the previous block, a time step, a state route, a transactions, etc., etc., many more. And what we want to do in order to execute or verify one one block, you need to execute transactions in that block.
And to execute transactions in that block, you need some of the state that those transaction access. Currently, nodes have all the state. So they have everything that transaction could possibly need and that's the only way to do it with statelessness. You could have something that is called execution witness. Um the execution witness basically is contains the state that is only needed for executing that block and with the proof anchor to the state of the previous block.
So if you have execution witness embedded within the block, it can have the state that is needed and the proof of the previous block that that state is a valid state and then the block can execute that one block at a time with all the data is being there. they don't need any entire state to keep themsel um so what what what do I mean if when I say state I'm talking about basically accounts their balances um other parameters like nuances smart contracts like just the bite code of the smart contracts and also like what is stored in the smart contracts like storage of the smart contracts um and uh Currently the entire state is called in a structure called Merkel Patricia tree. If you are doing solidity or maybe some protocol development like that you might have heard about it. Um and it one of the properties of the structure is that any piece of quantity is provable relatively simply. Um downside of it is that in order you can already do this in a with the Merkel Patricia tree.
The problem is if you want to prove the entire block the proof of that entire block would be too big. Now it depends on the average size etc. There was some research technically you can already do it but the the proof size would be too big to be practical um in a real case because all these block execution proofs would have to be transmitted to every node in the network and that would be order or two orders of magnitude much bigger than what they are already doing. So it's just too big to be practical at the moment. So we need to make the proofs smaller.
So what is the solution? We use snarks. That's the direction we are going. Snark stands for succinct non-interactive argument of knowledge. Um it is related to the zero knowledge technology.
Um and what is the idea is that uh we need a new structure for the state and that is where the current ongoing research is going. uh we have one approach called worklet trees with the EIP spec like where it's speced out it's been already done couple of years uh a lot of research been done into this um it has one downside that was known from the beginning and that is that it's not a final state even since the beginning of working on the worklet trees it was known that there will be another transition after the worklet trees reason for that is being worklet trees are known to be non-quantum secure and we will know that eventually Usually we need quantum security. Uh but when the work was starting there were no good quantum secure alternatives and the work starts in the direction. Recently maybe in the last couple of months there is being uh research going towards something called unifi binary trees which are quite similar in many aspects to the worklet trees but they are quantum secure. So potentially we could go directly to unifi binary trees instead of going to worklet trees and then switching to something else we can go directly but the work being done on the unifi binary trees is significantly less.
So the research is not as far but if we can go directly to that one then we only need to change the state only one time. So there are some benefits on both direction and there's ongoing research of how both of them would work. Again, cryptographic primitives for the worklet trees are way more stable. While the quantum secure primitives are also like still in a research regardless of the statelessness context, they are also like relatively new and maybe not as robust and not as researched as others. So there is a website statelessness uh FYI you can scale the QR code more information about the ongoing work projects research etc on this topic.
Um okay so let's see what the stateless clients actually gives us. This is the image from before. So once we have a stateless clients what can they actually do? So they can basically just verify blocks. They cannot do anything else compared to light clients.
So they can do everything light clients can do which is not much but they can very effectively verify the blocks. And that's where the basically the main problem starts with stateless clients is they don't have a state. So they cannot do almost anything except verify blocks and they're very good for that. It's already good concept. Um but here is this is the portal network is a project that I'm working on within Ethereum foundation and that's where it comes to the story.
Everything about portal network take a bit of um question mark let's say at the moment because just yesterday was announced the big reorg within it foundation and the future of the portal network is a bit unclear at the moment. Um but if this doesn't happen to be a portal network it might change to some similar project or protocol that will serve similar purpose. Um so have that in mind. Um okay so the portal network is basically protocol adjacent to the mainet or main like Ethereum blockchain data. It's another decentralized it provides decentralized access to the Ethereum state.
Uh what does that mean? Well imagine archive node that has all of the data in one machine that I said is around 10 terabytes. Imagine that being spread around the the peer-to-peer network similar to other peer-to-peer networks you know like for example torrent or something like that where every pit of content is provable belongs to the mainet data and you can access from other peers the content that you need about you don't need to store everything yourself but it's distributed on the network and you can fetch it from there so it basically decentralized light access to Ethereum states more like a decentralized archive node uh and the execution clients already started integrating uh it's part of the progress of EIP44s where basically currently as I said the full nodes they have the whole blockchain history so they don't have the state of the entire history of the Ethereum blockchain but they have all the blocks and all the transaction that ever happened so they can execute to those from the beginning if they if they really care about it um And the EIP44s is basically something called history expiry where execution clients want to delete some of that old content and not have it locally. And they integrate they're working towards integrating with the portal network to use portal network to get that information that they need it. So basically instead of every uh full node storing the entire um history of transactions, they're going to drop some of it and keep maybe 5% maybe 10% of the whole history of the blockchain and the rest like everybody if everybody keeps 5% or 10% there is enough redundancy in that peer-to-peer network that everybody has access to all the data that they need and they're working on integrating the the portal network in their clients.
Uh you can read more about portal network on that website. Um okay so what would potentially so this was the previous chart and now let's see with portalet client. So portal client would basically provide all of this functionality in a trustless and decentralized manner. Of course the the the check mark is a bit smaller with an asterisk because it's not going to be as efficient. When you have a peer-to-p peer network, you will be able to access all of this data and have it locally and verifiably uh without any trust exception, but it's going to be slower than if you have all of the data yourself.
Uh the portal network is designed such that you can decide how much you want to contribute. You can configure your client. I just want to contribute uh maybe 1 GBTE or 10 GB of data on my local hard disk or I just want to store 5% of the old state. No matter how big that is, it will keep growing. So you can configure how much you want to contribute to the network and you will have access to the entire state from other peers.
Of course it's going to be slower um than if you have everything locally. So that's why like the user experience if you want to use it for like a userf facing feature or something like that it's not going to be as great but you can have access to all of it. And because of that latency, it's not good for uh building blocks for example or some other uh latency uh dependent operations. The portal network wouldn't be a good solution for it. But everything that requires or allows either a small state or a bit longer latency, you can probably use it.
Um and that that is a solution. So if there is one takeaway from this entire talk, that would be that lightweight decentralized clients are on the way. They're being researched and in progress. And yeah, thank you. Any any questions?
[Applause]
Yep. There's a microphone.
Hello. I remember that last year at SP grade you um provide us three days meet up just the three day event for u for for for presentation presentation of this uh progress of portal network. I remember that I I also uh uh took part in that. But I wonder that besides the portal network and the currently unified binary tree, do you uh have any uh uh do you know any more progress on the data structure optimization of Ethereum data storage uh besides vero and uh and and the portal network and the and unified binary. Do do you know any other progress?
As far as I'm aware, there is no other um research or progress being done on the uh on the actual state and the structure how the state is stored. I'm not aware of any other research other than what you said work with trees, unify binary trees and portal network is not doesn't really change the structure. It just change how or introduces a new way of decentralizing and storing the state. So it doesn't really change the structure just introduce the more decentralized way of doing it. Uh there are other work on being done but sort of unrelated stuff on indexing maybe the data and uh distributing the the index data across the network etc.
So that for example even if you have access to the oldest state you cannot effectively answer the question give me all my previous transactions like you would need to go and iterate every block and try to find where are your transactions or you would have to index them before and then have access to it. So having access or index data also helps with like a user experience for like a most basic queries like if you have a wallet you want to know all your previous transactions and that's not very easy to do. You have to do some kind of like indexing. So there are projects being worked on doing that kind of stuff that is more like making the access to the data easier but there was no other project I know being researched that actually affects how the data is stored u on the Ethereum mainet as far as I know. Okay, thanks.
And I hope to know more uh research grant uh opportunities uh provided by either is in foundation of the portal network program uh to uh to boost the research of the uh data structure uh or data storage optimization further optimization and I hope to keep contact with you on the top port network discord. Okay.
Okay.
Any other questions? So um for the virtual trees that as you already just said actually um we have a problem with the postquantum security right so and you said like that there are just two ways for or the two research path for the state business um I just wonder like why didn't we try to find another postquantum secure commitment style for just achieving the statelessness. For example, we have like latest based postcontam commitments that we can prove the whole tree with uh to like prove the correctness of the you know execution and access but um I couldn't like any find reason uh to not trying to find the like latest space commitment. So I'm not actively working on this so I might not provide the best and correct answer but I will tell you to my understanding at least the unifi binary trees are designed such that um you can actually use any hash function inside the the string like it doesn't really matter what it is um as long as it has certain properties that we want like it's easy to prove it's easy to aggregate and make multiproofs etc etc um I think the most the the the the the one being considered the most at the moment at least by the researcher doing that is Poseidon. Um the risk of some of the new hashing function that might be better more performant is they didn't uh pass the scrutiny of time.
So a lot of them and even Padon is questionable is that they weren't basically they they weren't known for too long. the research community didn't have enough time to test them and verify that they are valid or try to break them and understand the guarantees and if you want to put it on the mainet and secure billions of dollars you need something that is very secure and trusted by the you know research community and everybody I think that's the main reason why there was no work on quantum security when the worker work started a couple of years ago because there was no good functions or they were not good enough that were trusted enough and the even the unifi binary approach is sort of like okay we cannot ship this on the mainet in the next one and a half two years and by that time we we are actively working with researchers to prove that Poseidon or whatever we are going for is actually the good like that's one of the one of the challenges of unifi binary trees not so much about how the structures look like but all the other like security properties of it etc etc how fast it's going to be how performant it is et so there are a lot of aspects to it and yes creating the proofs and how long it takes takes to create the proof is also a big factor of which kind of function you're going to use and I don't know all the properties of all different stuff. So the researchers are considering different functions. Um I I don't know exactly which one how exactly they work, what are the trade-offs between them. Um not really my core expertise but I know that how long certain functions are being known and how well they be tested by research community and accepted that they be trustworthy is a big factor they are considering.
Um, so yeah, that that's kind of my answer. Um, all right. Any any other questions? Thank you.
Automatic transcript — names and jargon may be misspelled.