# Erigon 3 a New Paradigm for Ethereum Clients by Mark Holt | Devcon SEA

- Channel: [Devcon](https://streameth.org/devcon)
- Date: 2025-10-09
- Duration: 26:10
- Topics: Science & Technology
- Watch: https://streameth.org/watch/yt-sMPe1Ae99aA
- YouTube: https://www.youtube.com/watch?v=sMPe1Ae99aA

## Description

Erigon 3 represents a step change for Ethereum clients:

* Modular client combining EL & CL
* Transaction Centric
* Deterministic storage model built to optimize EVM based chains
* Performs on commodity drives
* Sync model uses verifiable data replication and minimal re-execution
* Acts as block consumer and producer, RPC, or indexer
* Splits chain dissemination from chain distribution

This talk outlines the key features of Erigon 3 and explains how it will change Ethereum client  landscape.

Speaker(s): Mark Holt
Skill level: Intermediate
Track: Core Protocol
Keywords: Architecture, Data Availability, Scalability, modular

Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon
Learn more about devcon: https://www.devcon.org/
Learn more about ethereum: https://ethereum.org/ 

Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more.

Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. 
Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024.
Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

## Transcript

Uh. Okay, hold on a sec. I'll get the slides in the right place. Um yeah, hi. I'm Mark Holt. I'm work on the Aragon team. Um and basically what I'm going to talk about today is Aragon 3, which is the kind of latest incarnation of Aragon, which is just coming off out of alpha. So, we're expecting by the time we get to um Petra to have that as the client we would like everybody to use. So, the um purpose of this talk is to talk about Aragon and what it's doing and why it's slightly different from other clients. Um And what do I Oh. So, yeah. So, Aragon a a bit about us. Aragon is a um is a is a client team. Now, although we're principally used for execution, in fact, one of the things about later versions of Aragon 2 and Aragon 3 is actually it's a combined client. It basically does both CL and EL in one application. Um now, it's based on a a turbo turbo geth, which was a project which I'll talk a little bit about. Um but it's essentially built become mainly an arc uses an archive client, which is you see we're down at 2% of usage. Um and that's cuz basically we get used for um people who've got who've got an interest in the whole history, essentially. And historically, we've not been quite so good at at operating for people who are purely doing validators or are interested in the tip of the chain. Now, hopefully with Aragon 3, that will change a bit. Um So, what I'm going to talk about today is just an overview of what Aragon 3 is, the journey we've been through from TurboGeth, which is where we started from, to where we are now. Um I'll explain a little bit about the architecture of of Aragon 3 because I think it's important in um in understanding how we think it's different from other clients and and what what the implications of that are. And then I'll talk a little bit about the future of what of what we've got. Um so, I called this um this talk a Aragon 3 a paradigm shift for clients. And that's partly because I think within the Aragon team itself, we think about the client slightly different from other clients at the moment, especially in I I guess the Ethereum space in that we don't really think about consensus versus execution because we've got a combined client. Um what we think more about is what happens at the tip of the chain and effectively what happens after that because a lot of the technical work we've done is about the transition of data as it goes through the um life cycle from coming out of a mempool to getting into consensus and then, you know, what happens with all of the data afterwards as it gets into the um in in in into the archive space because I mean, for us, we never throw data away and that includes both CL and EL data. So, we're basically working out the how do we how do we store and distribute this stuff? So, we've kind of got a model that chain has an application that thinks about chain dissemination, which is, you know, the live chain and what happens when it operates in real time and how do we work that effectively. And then there's a whole process of once you've basically got the chain data, how do you then distribute it efficiently to other nodes? Um and we kind of split that and we don't necessarily do that through the entire the um the same life cycle as everybody else. So, um I'll talk a bit to start off with about the Aragon journey cuz I I think it's interesting. It's also interesting in terms of the last conversation we had, which was about client diversity. Um I think one of the interesting things about client diversity in the context of Aragon is it's not just for the live chain. I mean, one of the things about client diversity is Aragon's had a quite a long journey to get from where it started from to where it's ended up now. And that's effectively been funded out of the community because it accepted there wasn't a one size fits all. And because there's not a one size fits all, if you believe enough in about something, you'll get the backing to keep going, although there are several other people doing the same thing. And that can lead to different outcomes. Um so, if we if we talk about the where do we start from? So, um I think Alexey, who's actually in the audience here, was was talking in Devcon 4, which is and and introduced TurboGeth, which was a change to the way geth stores data. So, that was in, you know, Devcon 4. We're now several iterations through the um through the journey and we've we've probably finally come to the point where we're delivering on the original vision. Um and I think that's a that's another thing about the um the space that we're in is that actually, if you look at Aragon, it's had the space to kind of learn over seven or eight years to get to a point where it's actually delivering a product without the kind of pressure of you get you got get time to market. And if you haven't if you failed within the the the first couple of years, you kind of you've got go off and do something else. So, there's there's there's longevity in this process. Now, I think if you look in TurboGeth, that what it was doing was looking at the data storage and saying, if we have a look at the way data is stored in and do an analysis of it in in the um in the clients, we can do better with a storage model, effectively. And I think the the TurboGeth piece of the journey was about essentially, how do we make storage more optimal? And if if you look at the numbers, um the the the lesson there was you look at the majority client and we we with the analysis that we did and and the work, we got down to a fifth of the storage model. So, you get to a situation where the effort was worth it to get something slightly different, essentially. Now, I I think what that does is says, right, there's a product here and people start using it because there's a there's a benefit for people because as as the chain grows, this size is important. So, having a client actually cares about this is an important process. So, that's what got the whole thing kicked off. And then, essentially, what we ended up with there is an ongoing development of Aragon, which took that original idea and then pushed it to its um next conclusion. And the the next stage of this journey was the realization that actually, if you start caring about the data and the data going to a database, you start thinking about what the data space is doing and asking yourself the question, can you optimize that? And the other thing that you begin to do when when you look at that is say is see, even in the data storage space, there's a pipeline of um different things that you need to do in the storage process. And if you if you want to do different things, you need a slightly more adaptable client. Um so, essentially, what you ended up in Aragon 2 was this concept of a stage loop where the stages of storage are done independently from each other. And if you want to do slight something slightly different in the process, you can insert something into that model. So, you get a bit of flexibility. And if you look at quite a lot of the Aragon forks that people have done, quite a lot of the adaptations in those are adding stages into the stage loop. Um The other thing that Aragon 2 did is is realize that blockchain data has actually got a series of stages to it. And especially now the chains get finalized quite quickly, you've got stuff that you need to do in a read-write database, which is has got a a a certain set of criteria that you need to meet. But after a certain amount of time, the data just doesn't ever change, effectively. So, you've got this model where you've got live data, but you can transition that live that data through a freezing process into essentially a different storage model. And the interesting thing about that different storage model is it's a lot easier to distribute. You don't you don't need to take that data and push it through the whole execution process to get it onto a disk for a machine. You can just transfer it by a some other medium. And essentially, what Aragon 2 does, it says, well, that frozen data, you can just put it into a torrent and pass it around using um torrent information and you get the same data on the on the same disk. Now, Aragon 2, for anybody who's used it, um will see that process of you load the files when it starts. But then it's got a downside, which is you still have to take those files and execute them. And that's because in the Aragon 2 version of the system, the state of the of of the chain was still had to be computed cuz we had no way of dealing with the how do you store the state and redistribute it, essentially. So, um what what then happened is we moved into okay, fine. If we're going to actually make this distribution model work, we need to actually do it for everything. And it was quite a large effort, I would say, which is the Aragon 2 to Aragon 3, which has been, I guess, ongoing since the merge. I think that that that process. Um but But that was doing was kind of saying, "Well, actually the model of distribution that you've got for um transaction data and transaction history, you can also apply to state." So, um if you look at the picture here for um Aragon 3, um the what I'm trying to show in this is that you can take the Turbo Geth model where you're dividing the data out, you can take the Aragon 2 piece of the model where you're actually taking things from the data store and pushing it onto disk, and then basically once you've got that data out of the store onto disk in a form that you can distribute, you can distribute the whole chain independently of the um of of the consensus mechanism. Because what you're actually delivering at that point is just the frozen chain. And I think the way that I kind of conceptualized that is we've now got the same process as you do if you're basically have got a um you know, we're very used to when you compile code, you compile it and then you distribute it in its compiled form. You don't constantly recompile it cuz once you've once you've got a the compiled data, you can deliver it and link it in a variety of ways. So, it's it means that you can think about, you know, what we're doing in Aragon is we're taking the blockchain and we're compiling it into compiling it into a redistributable form. Um now, that probably wasn't the original intention, but it's effectively where we've ended up. Um so, I'm going to talk a bit more about how that works by giving you um a bit of insight into the way that the um Aragon architecture works, so you can see what we're actually doing inside of the model. So, essentially a lot of um what Aragon does and a lot of what the team work on is this kind of process where um we we think about what the EVM's doing and all of the the peer-to-peer network, but actually a lot of what what we're coding is this process of taking the um the blocks that we've got and the and the EVM execution and pushing them into the machine's page file so they can have be efficiently stored on disk. Um and then essentially they will either go into the database if they're read-write data, or they get pushed into files and then read. And a lot of the basic performance that you get out of Aragon is essentially the uh the layer we've got which actually manages how to use the machine's page cache. Um which is why if if you run Aragon, you'll see you may have a relatively small process, but if you've got a machine with memory on it, we will take all the memory that the operating system will give to us in its page cache. Now, it's to some extent that's quite alarming for users when they first see it. I mean, the the good news is that if somebody if something else wants wants the memory because it's operating system managed, it does get it back again. It's just um we'll basically eat eat as much page cache as you like, but it's it's useful to think about I think the the the the chain being processed and then put into a set of pages that put onto disks that are are that are just managed. Um and effectively the data goes into a database if it's read-write and and as it as it essentially um ages, it'll get pushed into um a set of read-only files. Now, essentially those read-only files are used by anybody who's using us as an RPC provider because the majority of the historical data will actually get read, especially in Aragon 3, out of these um out of these files. Um and the other thing that's going on inside an operating Aragon process is we've got this bottom line where we're basically continually pulling stuff out of the database, pushing it into these flat files, and then reconstituting the data. Um so, when we're thinking about the performance of what we're doing, we're thinking about essentially how do you um how do you balance this that you've got the EVM running at the tip of the chain, and effectively we've got this process of pulling stuff out of the database and pushing it onto the disk, and we need to coordinate all of that in the at the same time. Um The other thing we're doing, which is what this picture is is doing, is we're we're effectively in that process, the data that's pushed through the operating system's page files are also formed in this in a um in a flat state, which basically then we distribute across the network. And the reason I've drawn this picture like this is effectively what you've got on every Aragon node that it's running is really you can think of it as a replicatable page system where essentially um we're the data that's in the database between two machines is going to be different because it's it's it it's been loaded and and um and processed depending on on what the database on that machine is doing. But the files that we've got um that we store on the disk in in our background process are exactly the exactly the same on every machine in the network essentially. So, effectively what we're producing on N machines is completely deterministic. It will always be the same set of files. And it's that feature that allows us to basically um if you like, push them through the network as a set of pages. So, what what what we're doing when we're distributing stuff from Aragon is we're essentially distributing a verifiable binary version of the chain from one machine to another essentially. And what this diagram is is showing is you you've got a set of pages and they're all hash verified. Uh now, what what are the implications of all that um uh process? Well, there basically what you end up with is very quick sync performance effectively. So, the big difference between Aragon 2 and Aragon 3 is the amount of time it takes to sync something. And this is um you see this particularly on uh large uh chains. So, um basically my role in Aragon team, I mainly work on Polygon rather than rather than Ethereum, which is essentially a um I mean, on on Aragon 2 it's about 8 TB uh of data that needs to be moved. And if you sync it, you can literally wait weeks, literally. So, um with Aragon 3, effectively that time goes down to only a couple of days, which seems like a long time, but actually in that chain it's not a lot. And effectively for Ethereum, you're getting down to be able to sync within a few hours essentially. So, and just talking as a developer using a chain that actually completely changes the way you're thinking because the chain's there, you know, comparatively almost immediately. Um the other thing about our s- background syncing process, it scales with network bandwidth. So, um I'm I'll go back to Polygon an example. So, basically if you if you stick an an Aragon 3 node in a data center that's got enough network bandwidth, this um this 4 TB um chain that you're trying to sync happens in a couple of hours effectively. Now, and that's essentially because we're not reprocessing everything all the time. We're We're essentially just copying everything across the network, and as long as you've got enough network bandwidth, it happens quickly. Um and what this diagram is showing you a bit about is why does that happen? Well, if you think about the distribution effectively across the peer-to-peer network, you're you're effectively taking the data, pushing it into the network, and then every machine on the network is basically running a consensus interpreter, an execution interpreter, and then putting it onto the disk. We're basically doing a straight verified copy. And the thing is that you get exactly the same result in both paths. It's just the the second path is an awful lot quicker. Um now, it's also because each node pro- produces exactly the same files, effectively you know you can know between two nodes that you've got this straight data copy, but effectively because of the hashing involved, you also know that this is a copy of the chain effectively. Um so, you know, that that that's why we we think that's a change. Um and then just quickly at the end of it, I want to talk a little bit about the future. So, um the other thing that that does is because we've now got essentially a thing that's distributing data for our application, we're beginning to think about Aragon as as splitting our client into two things. One is a database, which is effectively running a DLT data store, which we're we can separate as a separate thing, and then a set of components that actually operate the chain. Now, I think what's important for us as a team is over the next year or so, we're probably going to spend a lot of time re- dealing with data and a lot more time about dealing with this front end of effectively how does how does the chain operate. Um and I think the other thing we can start thinking about is essentially we've got this model for distributing a compiled form across the across the network. How do we push that and start thinking about having sparse clients that don't bother downloading everything all the time cuz really if you've got a fast network, you can actually afford to not do anything until you need it and then load pages on demand. And that that point that distributed version of Erigon looks like the very local version of it where effectively we've got a two-stage page cache effectively operating the chain. And I think that with that I'm out of time and I'm done. Thank you. So two of these questions I'm going to combine which is how do you compare in terms of speed and cost efficiency in storage compared to the other execution layer clients? Um, depends which client. So something like So we're we're I think we're probably still and don't quote me, we're about a fifth of the storage for Geth on on a reasonable size size chain. The others I'm actually not quite sure about. Are there other pros and cons of running Erigon versus any other execution layer clients? Um, yeah, so Erigon is good in in storage performance and it's RPC layer. It's not so good if you're doing things like validation because we simply haven't worked on that as a we've optimized the storage not the interactions. But I think that is less than true with Erigon 3 effectively. Great, thank you. How much bandwidth overhead does Erigon introduce to push verified data across the network? Um, as much as you want to give it. So basically what you can do when you run an Erigon node is you can tell that how much bandwidth it wants to give away and it basically if you don't want to have that turned on, it doesn't take any it doesn't take any overhead if you if you're if you're conversely a seeding node, you know, it'll it'll run as much bandwidth as you want to give it basically. Got a lot of questions coming in at the last hour. Is Erigon 3 going to make the support of other EVMs more easily? What about L2s? Um, yeah, I think it is. I mean one of the things that we're explicitly doing with Erigon is doing L2 support. I think the the big difference for us in in 2025 from 2024 is we'll spend a lot more resources in in in actually actively working with L2 chains to get an L2 version of of Erigon and one of the reasons for componentization is so that we can we can have a core Erigon that does the database thing and then as we work with more chains, we're simply building components cuz what we want to do is end up with not having endless forks of Erigon in order to support other chains but have an ecosystem that means that you can plug in and extend it essentially which is something that we haven't been able to do while Are slot storage of a contract in the same page file? Our what? Top question. Are slot storage of a contract in the same page file? Um, they will be if they're close enough together in storage basically. So it it the question is it depends where they're stored. Will it be possible to query underlying database direct DB with directly without JSON RPC? Um, it it will do but it's not at the moment basically. So I mean when we talk about Erigon DB, what we're talking about is can we give it its own data access layer that effectively runs alongside of the RPC layer. So it actually operates as a as a DLT database basically. But that's where it becomes a product in its own right and not part of Erigon as a client. What is the biggest performance bottleneck in Erigon right now? Um, it's disk IO. Sorry? It's disk IO. All right, the the amount of time it takes to write to a disk is the thing that actually So it doesn't matter what you do to optimize it, it's how much it time takes to get stuff onto and off of the disk um, basically. So I think this question is similar. For transaction processing throughout aka transact transactions per second, is disk access access the main bottleneck? Yeah, yeah, absolutely. Is the separate RPC daemon process staying? Yes. Did you consider changing MDBX? It's sanctioned software by a fanatic. Um, we did but we've got a lot of Russians in the team so they don't actually have quite the same view of it. Do you support Verkle? Uh, we we have a Verkle implementation, yes. Uh, is it is it is it going to get deployed? Well, that's I think not not the same question. How do you optimize the database so that old and new data are not stored right next to each other? I don't know the answer to that question. Uh, wouldn't sparse clients contribute to storage fragmentation or loss? Um, I don't think I don't think they would do actually. I think they're a natural extension of what we've done cuz the question becomes, you know, if you've got a lot of clients and you've you're all storing a bit of the network essentially, you won't necessarily lose it. I think I think the the question there is more I think part of where we're going to is you've got to answer the question is if we want an internet scale scale blockchain, you can't have all the data on all the clients all the time. You have to get to a situation where you've got some kind of sparse storage basically and then it becomes a how many people have got how much of the chain. Cool. Thank you so much. We are out of time for questions but if you want to ask these questions to him afterwards, find him afterwards. Can we give one more big round of applause?
