New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Securing Ethereum Pectra before mainnet - Zigtur | Spearbit & Cantina

ETH Belgrade CommunityTue, Oct 7, 2025, 12:00 AM

Securing Ethereum Pectra before mainnet - Zigtur | Spearbit & Cantina

Transcript

Um so hi everyone uh today I'm going to speak about how we have been shipping ethereum pectra to mainet especially so what we are going to talk about is ethereum architecture first because we are all building on ethereum then um we will see the scope of ethereum pectra which was a big upgra upgrade we are going next with the test nets. Why do they exist? We will see that on the security view. We will go through the security competition that occurred for securing spectra and then we will conclude. So he said all that uh we don't really care.

So Ethereum is like divided into two software layers. The execution layer that you all know about you are user of it and uh the consensus layer. So as I said execution layer is users and apps. This is where smart contracts are and there is the consensus layer where all validators are interacting with each other. So an interesting point is that there is communications between the execution layer and the consensus layer.

When you want to create a validator, you will just have to interact with the deposit contract. Um, this is where the well-known 32 is uh is deposited. It will emit an event that will be consumed by the consensus layer to create a validator. So what is Ethereum pectra about? It is a set of EIPs that redefine parts of the ECM protocol.

EIP are Ethereum improvement proposals and some of them were on the execution layer. You have been hearing about 7702 which transform your EOS into some sort of smart contract. Uh but there are more and there are ones that are really interesting that are just between the execution layer and the consensus layer. Um and a lot of bugs have been found in there. So Ethereum is divided into multiple consensus clients and multiple execution clients.

So to make the upgrade you have to implement all the IPs into all the consensus clients and so doing that you can be pretty sure that there will be bugs because it's a huge task to do. So first one we have been upgrading the Oleski uh test net and there was an incident instantly. So what was wrong? There was a deposit event just as I explained before which uh the deposit contract is address 0x4232. The thing is that guess nezomine and bezu were not listening to the correct deposit contract because it was not set.

So it was just defaulting to listening to address zero. So when a deposit occurred in 0x4242 rest which was correctly configured was saying hey there is one deposit but guess nezomine and bezu we're just saying there is no deposit in this one. So when you have that this mean that at the consensus level you will just handle a deposit on one side but not handle it on the other side. So you will have a chain split where validators are not having the same state. So it was fixed pretty quickly.

The issue was not that big. It's just a JSON file that was wrong. So nothing to be scared about. Uh but what is interesting on this one is that the issue happening at the execution layer happens to have led to a lot of impacts on the consensus clients. So memory consumption that was not expected uh and things like that.

So we have seen a lot of fixes coming on Lighthouse, Pris Lordar even if the issue was not coming from this specific code bases and then we went through the CIA upgrade an incident once again. This one was interesting like it was just creating empty blocks. So the blocks the chain was continuing processing but empty blocks and this is really specific to Saporia like the mainet. So we we have to keep in mind that we are all building for mainet we don't really care about test net right we are looking at mainet building for it. So if something is not the same on test net then it may go wrong because it will go good on mainet but not on test net.

So for the EIP6110 which redefineses how deposit events are listened to um we were expecting only one type of event coming from the deposit contract. The thing is that is that the sealia contract is different different and is emitting a transfer event. So when the EIP6110 has been implemented, it was not expecting this event and so it has led to an error when uh processing this type of event that is not expected. So when an error was triggered in the passing of the event then the block created was just empty. Uh how was it fixed?

It was just that on each log that we are uh listening we just check that it is a deposit event and not a transfer one. If it's a transfer one we just ignore it. So test net are here to spot issues that's for sure. Like if you are scared I mean we we have seen a lot of fud about that like all SK event uh the the simple event then mainet is dead. No, no, no, no.

I'm the security guy. If you are scared, I'm not. This is just bad luck. And in the same time, there was uh a security competition happening on Canina offering up to 2 million USDC. The results, three high, three medium, and 17 low severity issues were found.

The first bug uh is on lighthouse which represent approximately 40% of the validators. It was found out by Alex Filipov and it's so the specs says that at the end of an epoch which is a collection of blocks you have to process the pending deposits process the pending consolidations and then update the balance of the validators which gives the graph call we can see on the right. But the thing that lighthouse implemented was that in the process pending consolidation, it was doing a call to the process effective balance. So it was updating the balance. And this may not have been an issue, but it was like according to Alex Filipov, updating the effective balance twice will not always produce the same result as up updating it once because of eststerosis.

Sois is just like you have to cross a given threshold to update the balance. So if you do if you add twice 0.5 is it will not be the same as adding once one is and this was the issue. So it was unling handled correctly but due to this specific historicis thing uh it was broken and then you can have a discrepancy in the balance which leads to computing a different state which leads to breaking uh having a chain split which breaks finalization at the end. Then uh there was a bug in prism that I found uh this one was on process pending deposits.

So basically when creating a validator if it's the f the first time you you you do a deposit for this validator you will have to provide a signature and this signature must be valid to create the validator then if it's the second third or anything deposit for this validator we don't care about the signature but pris was doing an aggregation ation of signature because you can do that instead of checking each signature in a block you can just aggregate all signature and ver verify them at once. The thing is that if you produce two deposits in the same block for a validator that was not known the second deposit you should not check the the signature of it but pris was checking it due to this optimization. So basically it gives the the this um this thing like first deposit valid signature it's a success second deposit no signature it should be a success for all nodes but prism was failing because the signature was invalid so same thing again you compute a different state because of the balance the deposit is just ignored and then chain split breaking the finalization. So what we can get back from this issues is that every time you just do not follow the spec for optimization purpose maybe for other things but you are not following the spec. So this is the big issue and how to avoid issues just follow the spec and an open question I don't have the answer to that is optimization an enemy of security uh but yeah we can follow the spec um this bug found by NDO was about elliptic curves the spec says that if any input um for for the the pairing contract is the infinity point then the pairing result will be one.

What was implemented is that if a point with infinity was found then we set uh we will set verified to true because that's what the spec says right but the thing is that when we look at the incident response uh report it means two things it was not clear clear enough like the pair make me maybe ignore but the multi pairing should be computed because you can provide a lot of points and basically nezoline was implementing the second if one of all the points is the infinity point then I'm just saying it's okay it's a success the others were computing the rest of the points and ignoring the one with the the infinity point which is really different so the the nezamine was fixed uh obviously but the spec Also, this specific note that was not clear enough was just removed from the specifications. So, we should follow the spec, but not always. As a summary, uh there was two incidents that happened on petrol upgrade. Uh these were test nets. This is a for uh then the security audits and competitions found a lot of security issues, but nothing critical.

No f no phones were at risk. Uh the biggest risk uh the biggest impact found was chain splits that break final finalization. It's it's bad but not that much. So what we should learn uh get back from it is that writing specification is mandatory no matter what especially for Ethereum when you have multiple um clients implementing it. You should follow the spec um especially as I said with multiple implementation you should also clarify your spec when it is not crystal clear.

I mean by that that the nezamine team should have asked the question and do not forget that bugs are here by default. Thank you. Do we have any questions from the audience? question about the contest. Uh when you were checking the code, were you checking all clients or only one implementation?

Uh so I did not have that much time and uh like to get the I severity issue or a medium. Medium was like splitting 5% of the chain and I was splitting more than 33% of the chain. And so when you look at the clients, um, Prism was more than around 30% of the chain at this time. So you know that if you find a bug in this one that makes it split, it's at least a medium or I if it's more than 33%. So I focused especially on the big uh code bases that are used like the ones that have the most share invalidators uh which and then taking the specifications and seeing anything that is wrong.

Um did that answer your question?

Yeah.

We have one more question over here. Hey, so about uh holism meltdown uh I I remember I think Lighthouse had seven better releases to try to fix it. So my question is do you know whether there was some processes changes that in the future if something like that happens that will be that recovery will be um uh more uh smooth smoother

about which one especially

uh I mean I'm just giving an example lighthouse like has seven releases about other clients I don't know but that's the one I followed so do you know whether are there any lessons learned is there any process now to make it make the recovery process any smoother for example people are always talking about like social consensus but as a for example operator you don't have any you didn't have any endpoint for example to choose your your chain for example and stuff stuff like that so

yeah uh yeah so this is difficult you you can always test there is the you can deploy a local test net to use it but when it's in production social consensus is the right thing at the end because we are all humans and if the for example for the the Oleski incident the chain finalized if I remember correctly so it finalized but with the incorrect state uh because it was guess nezmine and bezu they are they add more than 67% of the chain so it finalized but with the incorrect state and so reverting that means that all validators should agree to get back.

Yeah. Yeah. But but I'm saying like you don't as a user you don't have you don't have any tools to force social consensus other than just downloading the what's the latest client and like hoping that it will it will work.

Yeah. Yeah. Yeah. For sure. Yeah.

Yeah. It's always the team that has the control sort of. Yeah.

Yeah. I would I would have liked that if it were like uh the the we learned something about it and then I don't know how but that some some tooling is developed so like users can choose uh their path their their path forward.

Okay. Yeah. I did not about that. So sorry.

I hope that did answer the question. Okay. Wow. I I I I love security community like I I I really do like you guys are engaging. I I love it.

I love it. Rahul, I hope you'll be a gentleman.

Thank you for the talk.

Uh generally curious. I'm not familiar with consensus clients at all like working primarily with solidity. Uh how easy for you is to spot the bug in this code base? Like I prefer always play with things. Yeah.

like uh spin it up locally, writing tests, trying breaking things. So what is your procedure generally is it if it's not secret like can you play with that break it like locally see what happens and so on.

Yeah. So for this specific contest um it was like I mean is a set of rules that is well defined so anything that doesn't comply with these rules is a bug and may be exploitable. So that's how I found some bugs in there. But most of the time, yeah, I need to really understand all the flow uh of the codebase. So where are the user inputs coming from?

Where are they going? Uh for more causes things when there are multiple validators, what happens if one validator is malicious but not the others? Um, and yeah, you it it's just following the flow. That that's the thing that I'm doing. Um, sort of follow the money approach.

Uh, I know that some security researchers that are really good are just going through all the files and having a a mind map of it, but I don't. I need drawings to really follow the flow. Yeah.

Thank you. Hey, thanks for the talk. So I'm wondering so the um the contest was done after the test net incidents, right?

Uh actually the seleia events uh incidents happened during the contest.

It happened during the contest. So yeah

my question is um for the bugs that you found in uh why were they not catched during the test net launch? uh be I so for the prism one which is the one I know the most is that you need to do two deposits with a validator that does not exist yet in the same block which doesn't occur in the wild it occurs only if you are malicious uh or really bad luck uh but for the light was the light house one it should have been spotted I think before anything But uh yeah it's maybe sometimes it's too complex with the sterosis for example thing it makes it so complex that you can't cover every edge cases. Yeah it makes sense. Thank you. Um do we have more questions?

Yeah, thanks. It looks impressive. Um, I just wonder after you have found all these bugs, did the developers include uh did they introduce more tests to to discover these bugs in the future or they have to hire you again? Yeah. Uh, so I mean you can write a test but if you Yeah.

There there are so much edge cases. We write a test to prove it as security researcher. Uh but yeah thinking about it when you are a developer yourself I I think you have to have a security experience at some point and thinking yeah be paranoid.

Yeah but they have serum tests right? They have a test suite. Yeah, they they they have a lot of tests which are really high quality but yeah it's really difficult and especially when you put it the logic may be good but in the context of having multiple blocks which makes it a little bit different then you you would have to just simulate the chain and every edge case which makes it really really hard. Yeah. probably more like ethical question like is it really secure because like if everybody chases the money right and everybody knows this client is most popular I will get probably rewards there and nobody will look into other clients like that much.

Yeah. So is it really uh fair to share this like grade those vulnerabilities on current uh ratio of clients right because it can change in future that lighthouse prisma prism you mentioned yeah will be not the leading one but others probably and there like there is a bug some there yeah and something can happen but nobody look into that because it was like 5% and everybody told I will not look into that

okay yeah

it's more like philosophical question

yeah Um yeah yeah yeah security on the philosophy like I think that white hats may become black hats but the opposite is not true. Um so yeah uh yeah like what is really the question there it's not yeah sorry

uh sorry to be intricate probably so the incentive for security researchers uh like impact

yeah you mentioned that um Prisma has the greatest impact because of the share like which how much you can split

yeah yeah this is a thing with the bug bounty on ethereum it is based on the impact indeed.

Yeah. Yeah. But generally uh probably more security researchers will focus on those like high incentive Yeah. things but we'll lose uh will probably not look that much at like low incentive part. Yeah.

Yeah.

This was my question.

Yeah. Okay.

How secure is is it really? because for me it looks like it's pretty simple bugs. I always thought it's like uh cosmic things but it looks like pretty simple and I also was wondering about tests. Yeah.

So it's like unit testing not like end to end probably.

Yeah. And therefore I have my doubts right is it really secure because nobody will look at other clients which have small share uh and small usage and there are no incentive to look for errors there.

Yeah.

Thank you.

Yeah. But I think that most of the time you like if a bug happens in TEU for example which has only maybe 10% of the chairs it it doesn't have the same impact. So at the end you will put money to protect what is the most important thing uh and everything that is next to it is less critical and yeah we it's just put the money where the risk is. Yeah but yeah for sure there are there could be bugs everywhere in code bases that are not that much used.

Does that answer your question? Yeah,

we have time for one more question from the audience and I got to say you guys have been amazing. Amazing. Especially this session like I don't think we had this big Q&A and this much curiosity like in the history I think. So like I I love to see this.

Do we have any more questions? Of course we do. Okay. Oh yeah, that's the spirit. That's a maybe a spicy one, but can you comment on um your sentiment of should there be a follow-up audit contest on this update or um how safe do you feel personally after doing this contest?

And um maybe what can we do in the future to um like we what can we do in the future maybe to um uh let let's stick with the first question. Yeah. Okay. So what what's your your sentiment there? Is it safe?

Yeah. I mean most of my money is on it so I trust it. uh and I think that uh even if there is a security issue that could break something we it has happened in the past it's social consensus will be the norm at the end um no matter what the code may say we we sometimes say code is low which is not true because what we want is consensus between all the community and if the code has something that is exploitable we may roll back even if it's only if it's a big issue. Um but yeah, do do I feel confident? Pretty much.

Yeah. What we could do uh is I know that uh the SAM bug bounty is up to 250k which is pretty good but not sufficient compared to how much of assets it protects it handles. So that's uh something it should yeah 10x minimum.

Automatic transcript — names and jargon may be misspelled.