New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Beyond the Surface: The Hidden Benefits of Client Diversity by Daniel Lehrner | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

When people discuss client diversity, they often focus on technical benefits like resilience to DoS attacks or preventing catastrophic errors like finalizing an invalid chain. But there's so much more to it! Client diversity not only spreads the responsibility of maintaining the blockchain across multiple teams but also brings fresh ideas and perspectives into the mix. In this talk, I aim to explore all the hidden benefits of embracing client diversity. Speaker(s): Daniel Lehrner Skill level: Beginner Track: Core Protocol Keywords: Core Protocol, Staking, Home staking, diversity, client Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

[Music] my name is uh Daniel lner I'm a b developer at uh consensus and I'm going to talk about the benefits of client diversity so first of all let's start with who here knows what client diversity is can you raise your hand if you if you know okay I think almost everybody so for the people that do not know in ethereum to to use ethereum we have different software programs especially since the merge we need two pieces of software we have the execution layer which is basically the evm or whenever you call an RPC endpoint normally like gut balance if call and we have the consensus layer which is the proof of stake part and for each of those two pieces of software we have a lot of different implementations on the execution layer side we have GF which is written in go NE mind in net Basu which I'm working on in Java Aragon and go ref in Rust and the ium CH in typescript and the same on the consensus layer we have prism which is in go Lighthouse rust take Java nimos is written in Nim which is like a variant of um C or C++ no I'm not sure we have load star in typescript and we have grandine um in terms of distributions on the execution layer site historically GF is what we call a majority client this means the majority of users are using it um on the consensus layer side the distribution is more even mainly because they have all started at the same time they have all started more as it emerge um so that's a bit better so in the end why are we doing this so what are the benefits of doing the same piece of software several times and I'm going to use use this Iceberg metaphor here basically starting at the the tip of the iceberg which I think might be the most obvious answers to this and then we are going down to the less uh obvious benefits so the first one is stability that's the very top here we have an example where vitalic himself wrote on Reddit on the 24th of November 2016 there's a consensus flaw in G we have identified the problem and now we are in the process of testing a fix for a release so basically what happened is there was just a bug in G and at the time of writing this there was already a second client which um which was parity and basically what happened is the two clients processed the same chain and at some point because of the bug they they forked off and we had two chains so after a while of debugging they needed to find out okay which is the correct one and basically fix it so what what happened here is that even though GF which was at the time the majority client had a bug the chain never went down because there was par was still there basically to continue in the meantime parity contined to correct fork and once GF was fixed um GF rejoined this Fork so here is is basically I think maybe the the most obvious one if you write software more than once every every piece of software can have a bug but the two versions of the same protocol have exactly the same bug is rather unlikely so we can um basically make sure that one of the versions will work normally one one at a time then the other thing is correctness and error detection how do we even know that something in the protocol is wrong so I mean there are very obvious errors in software that's normally when the software is crashing but what if the software just has a bug and it calculate something incorrectly here we have an example from the 15th of August 2010 long before ethereum has existed this is the Bitcoin talk Forum there is a user that says the value out in this block is quite strange and then here at the bottom it says 92 b233 m720 368 Bitcoin basically what happened is there was a buck in the Bitcoin software and somebody triggered the BK we don't know if it was on purpose or not twice and created 180 billion Bitcoin even for people that do not know a lot of Bitcoin they most probably know that there is a very hard limit of 21 million Bitcoin it's 180 billion obviously is way too much and why I chose this example is Bitcoin traditionally chooses only to run one client normally what happened here is that this client had just a bug It produced in this case too much Bitcoin but the chain still continued to run there was nothing obviously wrong with the chain just that this one user was really looking at the transaction was looking at the output and that's how the error was detected in the first place this in the example before from ethereum was very different in ethereum two pieces of the same software ran and then we could see that they diverge so it was very very obvious that something is wrong here it was really by luck that somebody checked this and then at the end at this time satosi was still around they fixed the buck basically rolled the chain back um there were like 50 blocks afterwards produced they they were afterwards invalid um and then they were to head Val chain again then the next point is not just avoiding errors but especially avoiding catastrophic errors here um 6 months before the merge DKR Feist one of the ethereum foundation researchers wrote an article that says ethereum merge run the majority client at your own Peril so why did he write this um before improve of work in the first example that we saw when one of the clients had a bug it in the end produced an invalid Fork but you could always go back with the merge and with the change to proof of stake we have a New Concept which is called finalization finalization means at some point normally after what we call two Epoch which is around 12.8 minutes the chain finalizes if it has more than 2/3 of the votes and you cannot go back anymore this is a very very nice property because you know for sure if your transaction was in a block that has finalized it will never revert again but there is one problem with this I I don't want to go into too much details because it's basically not the goal of this talk but in the finalization process if a client a majority client who has more than two3 of the notes or of the stake has a bug it will produce the invalid chain a and it it will attest which is basically vote for the two blocks these are the two error uh the two red errors even if the bug is detected because we have other clients who will let us know that and the bug is fixed all the clients that were on chain a cannot go back to chain B anymore because in order to do that they would need to vote here for what is marked as the X and this is not allowed by the protocol because they basically jump over one of their own votes and this is what we call a slash offense slash offense means the validator does something that can cause the chain to not find consensus so we are not sure which of the two chains is correct and this is what could happen here okay so this this this was a very very serious problem and dunrod wrote this like half a year before the merge so the most logical outcome most probably would be that people would be aware of this and people and validators would make sure that no client had more than two3 at least that's what we thought the reality looked very very different so this is the decline distribution from beginning of this year and G had around 84% of the of the stake of all the notes so this means we had a ticking Time Bomb for one and a half years but basically one consensus bug in G would have triggered the outcome before and if this would have happened we would have destroyed billions of dollars of value in Eve in order to revert to the correct chain okay so now we had this ticking Time Bomb what what what should people do how how I I I'm one of the developers of a minority client so we tried our best to to make for software competitive but then unfortunately Murph is law hit and we had in B on the 6th of January a bu that basically took the note down and the second one is from nethermind which is the number two which was two weeks after So within two weeks the number two and the number three client had a problem so we had now the issue that GFF was had already this huge majority plus we had two problems surprisingly those two events led to the outcome that people finally were talking about the problem with the finalization of a super majority client here a big shout out to the to the a Staker Community to nixo to yor they they really educated people about this problem and stakers and validators finally realized okay that's that's a huge issue we need to do something and that's why right now it looks like this so G has gone down from 84 to 52 so it's below the 23s the issue basically this was outlined before it cannot happen anymore B and Ne mind have both increased a lot so now the situation looks much better um but what this also means is clant diversity is a choice the community has to choose it if not even though we have several implementations of the same software it will not happen um here is one example where this does not work if people are thinking about Bitcoin they must probably just think about one client in reality Bitcoin has 14 different clients some of them just on mobile because Bitcoin is more light client friendly others are full nodes but despite that 98.63% of all the nodes are running Bitcoin core so even though a lot of clients would be available the Bitcoin Community chooses to run only one which basically means they do not get the benefits from client diversity that they actually could have very easily um then here's another Point that's maybe already um a bit more hidden is to make the specification of the protocol easily accessible um when several teams across the globe work on the same piece of software you are basically forced to first start with a specification you're basically forced to first write down okay what actually are we going to do it and we have several resources in aerum for that we have the a magist forum which basically normally is where discussions are starting about changes we have the E process which is then the actual spec of the changes we have the etherum execution spec tests on the execution side which basically means all the changes that we are doing and all the protocol rules are available really as as tests that you can run so all the clients can run the same test and this is very clearly defined what is the expected behavior um another thing that we have that is less blind spots with client diversity you will never have no blind spots but at least you can reduce your blind spots um what I mean with this uh spe um is when we look where core developers are located we have um a big number in North America big number in Europe and a big number in Australia this means geographically more or less the distribution is not so bad but we have for example right now no team in Asia no team in Africa no team in in South America um if people were yesterday at Justin Drake's talk about the beam chain at least the situation should improve for Latin America and for Asia and we will have two new teams and basically the goal is that they can bring their own view their own life experience into the process and they can help us reduce our blind spots which developers from the from the Western World Maybe do not see so for example people who use stable coins in Latin America or Asia I think most cevs right now are are not in situations where they would really need this um then when you have different clients those clients can do different Improvement and experiments here I have just for every execution client one there there are a lot of them but we can quickly go over that NE mind has for example done a lot of performance work recently they have reached one gigz which is like our goal for layer twos in bestu we have a parallel transaction processing so right now on eum inet you can do parallel transaction processing without problems which is also something um that was long ago also for L2 Aragon has worked very hard to make archive nodes run on consumer Hardware especially the dis requirement went down from like 16 terab to 30 or4 ref has introduced execution extensions which basically allow you to modify the client without forking it and another example is eoff um if you don't know much about e if you're a smart contractor you will like it because it basically gets rid of the stack to deep errors and it finally allows us to safely increase the smart contract size limit limit eventually and this is mainly driven by the minority clients and then if you think okay I'm running majority client what benefit do I have the benefit is that the majority client eventually can copy the features for example here this is Peter the team lead of the go ethereum of the G team tweeting out that they will the XX the execution extension support from ref will also do it to G so because one of the minority clients introduced it finally the majority client gets it at well here is another slide which is a bit more abstract but you have a lower risk of protocol capture or failure there are if you have just one team there are different times of of of of attacks a government could do they maybe cannot detect the protocol but they could infiltrate the team or they can pressure somehow the team to do to push a change that uh they like a team can withhold updates something that the community maybe really wants to the team cannot force the community to run a specific software but the team can say okay this one update we do not like it we will not ship it and then the community will have a very hard time to get it and simply if you have one team they can just eventually abandon the client so then again you have a problem then we go to the last point which is in the end permissionless contributions uh to the protocol this is very abstract but in the end what this means if you have client diversity it means everybody can become a c of including everybody here in the room so if somebody is interested in this here in the slide we have the different repositories of the clients this is if you are interested in in programming um there also another things if you want to just keep up with the with the protocol on you can follow Tim or Christine Kim they they always share updates you can join the Discord server you can look at the EPS there are a lot of things to do so that's basically it so thank you for your attention and we will go through questions um if 4K never finalized a validator that voted in Fork a can later attest in Fork B when the fix is released can they do that without getting slashed yes so this means if the if the majority client is less than 23s um that is possible yes so the the the main the main thing is really no client should have more than two3 of of the staking votes um so we have some more time for questions we have more questions in the queue does anybody want to ask a Live question we have one over here oh there is one oh how do you get the numbers on different clients running different uh running nodes yeah okay so that's a good question um so in we have no reliable way of of of doing that um people do some sort of fingerprinting where they analyze different blocks that are proposed and then they they they can basically estimate which sale clients they are uh for the execution GLS we do not have that um there were recently proposals we have like one field when you propose a block where people can basically write random data there are proposals to basically write in which which El and which CL you run and then we have a best better estimations but the numbers that I showed before they mainly rely on self-reporting so the big staking providers in the end report the numbers and then there is like a huge chunk of I don't know one3 of the network we we do not know um but the for for the worst case assumption we just say it's the majority client and and still with this worst case assumption no client has more than two FS but yeah is we do not have exact numbers we only have approximation of the numbers thank you um could you Steelman the argument that Bitcoin puts forward for having one client and why it's wrong MH yeah so one client um in the end means um there is less politics involved maybe in discussing changes so in eum's case you basically have to convince different teams that they they want to change although Bitcoin is quite resistant to to change by itself um the other approach I don't I mean it's it's like what I mentioned before with with the capturing of the team um but yeah so so for me it's actually hard to steal man this um you you just avoid maybe I don't know certain errors by using more than two clients but yeah as I said for for me the benefits far outweigh the the down

Automatic transcript — names and jargon may be misspelled.