# ETHWarsaw 2023: Marius van der Wijden, Ethereum Foundation - Another Deep Dive into go-ethereum

- Speakers: [Marius van der Wijden](https://streameth.org/speakers/marius-van-der-wijden)
- Channel: [ETH Warsaw](https://streameth.org/eth-warsaw)
- Date: 2024-10-07
- Duration: 30:55
- Topics: main stage, day 1
- Watch: https://streameth.org/watch/yt-xCcK1r_iyrw
- YouTube: https://www.youtube.com/watch?v=xCcK1r_iyrw

## Description

Presentation - A talk by Marius van der Wijden from Ethereum Foundation. This presentation delves into go-ethereum.

Follow us for more updates: https://twitter.com/ETHWarsaw

## Transcript

presentation by Marius F about go ethereum thank you Marius give her a warm welcome yeah thank you for having me uh my name is Marius um I work at the etherum foundation for uh 3 and a half years now uh on on go ethereum mainly and um so this presentation it's going to be quite weird um but I I hope you bear with me and you enjoy the format um so I kind of up a bit um because um this presentation is meant to be more like a workshop um but I'm just going to try to do it anyway so how did this presentation came about um I gave a talk last year at eth Amsterdam where I went uh through the go ethereum code base and showed how the transaction flows um through different modules in go ethereum what we do there and um we started with submitting a transaction uh then we created a block out of it and in the end we ended on evm execution um today I want to do uh to trace another transaction but now we don't start at the E API which is the entry point where you send your transactions if you um send them via Json RPC um but today uh we want to start at the engine API so this is roughly how your node if you run an ethereum node looks like uh post merch you have two noes you have the consensus layer note and you have the execution layer Noe and um the execution the consensus layer node basically tells the execution L layer node what to do and and uh so there's a lot of new people coming in which is quite funny um and um this kind of glue between those two nodes is called the engine API and this is where we're going to start our journey today and um yeah and and and I'm I only know where to start I don't uh don't know where we end up so uh uh let's see let's just go with it so the engine API has they said is the glue between the consensus and the execution layer um for those who don't know probably not that many of you um G is an execution layer node um so it's responsible for executing the transactions and uh this this engine API has a few functions um that are standardized and uh the first function that we standardized was new payload um the consensus layer uses this function to tell the execution layer um of about a new block and uh we have to execute the block and say yes it's a good block no it's not a good block um but we don't update our current head so we don't uh set the fork Choice um that method is called Fork Choice updated uh with this the consensus layer tells the execution layer okay this block that I just gave you this is the new head um why is this split in two methods um because there could be the case where like five different blocks arrive at the consensus layer node um at the same time we need to make sure that all of those blocks are correct and then uh the the consensus layer will uh Aggregate attestations and choose which of those different blocks that are currently competing for the head is the best one right now and um this is what the the consensus layer note uh tells us in fog Choice updated um fog Choice updated has also a another small functionality uh optionally the the consensus layer uh asks us to build a block on top of the current head so the the consensus layer now tells us okay this is the head that you should have and please build a new block on top of this because we've been chosen to produce the new block and so this is what happens in in for updated and in get payload we can retrieve this pending block uh and then there are those are the big four methods that we need and there are three additional methods um that are just for convenience uh get payload bodies by hash and range um they uh return old blocks to the consensus layer this was uh this was implemented because uh um the consensus layer might not want to store um all of the old blocks um because uh the uh the uh the execution layer already stores all of the blocks um so we would need to store them twice we don't want this so we can we need this method uh exchange transition configuration uh makes sure that the execution layer and the consensus layer are both configured correctly and exchange cap abilities uh tells the consensus layer which of those functions are available and uh because we um we have to um we have to further the network we have to work on the network uh we have to change the engine API we have uh we have built in versioning into these uh functions and with uh exchange capabilities we can tell the consensus layer note hey I support these versions and I and then the consensus layer can can can say okay this this note doesn't support this version that I that I want so I'm going to show the user an error or something okay now to the fun part we're going to look into the code and actually see where all of this is happening um the engine API is uh is located in in in go ethereum is located uh in the e/ Catalyst package um it was called Catalyst before for the like this was way before any work on the merch was being done uh we we we started with this um proof of concept called Catalyst and um it's kind of like we never really moved the engine API out of this package or renamed the package so uh everything is still there and um what you can see here is uh when we do exchange capabilities uh we we send the consensus layer these methods so these are the methods that we support you can see fork Chas updated V1 till V3 get payload V1 V2 V3 and so on and um now our journey starts uh with uh oop sorry new payload sorry as I said this is the first time I'm doing this um okay so with new payload the consensus layer tells us this is a block that I received um please make sure that is well formed executed and store it but don't update the the your state yet and so uh we get this executable data which is basically the um so before this these functions are called uh the Json that we get from the from the consensus layer is is turned into these internal types well it's we're not going to look at that um but um yeah so that the that that happens before this method is called and so um the first thing we do is we turn this executable data into a block and so this happens here where we we have some transactions that are just bite slices and we have to turn them into transaction objects um and this is also where we do the the the first number of uh verifications on this data so for example after the merge um the extra data field in the block header is um has to be less than 32 um 32 bytes and so we check this here for example and um then we unmarshal stuff and um for Here We compute the transaction root um so the the header of the um the block header contains the root of the transaction uh of the transaction tree which means uh we take the transactions we build a Merk tree of them and this root of this Merc tree is put into the header and we need to make sure that the transaction that we have uh matches to this to this route and that this is what we uh what we do here oh actually this is for the withdrawals we do the same thing for the transactions here um and um then we construct our header and with this header uh we add the transactions to it and then we compute the hash and um we verify that the block hash is correct um this is just a basic verification that we have to do whenever we get a block we have to verify um okay if we construct if we take this data that we just got from the consensus layer and we create a block out of it is this a valid block does it compute to this uh to the correct block has and um so what we do in the engine API now is uh we we we we look in our local blockchain okay did we see this block already if we saw this block already and it was verified we can just say okay yeah we saw this we verified it in the past we will just return okay um if we have haven't seen it um then we need to check whether this is a known bad block have we seen this block and verified it and verified it to be bad before so was there something wrong with it and uh this is what we do in check invalid ancestor it checks whether this block um this block was ever marked uh invalid or if the parent of this block was marked inv valid or and uh any ancestors of this blocks block were marked inv valid so for example if I get a if I get a block and it is built upon another block that is invalid then this new blog is also invalid of course and um so after we check this we retrieve the parent block the parent block is just uh just the parent the the block that has the parent hash and the number in our number minus one so it's one previous and um then there's some stuff about the block time so um every block has to be be consecutive that means our block time has to be bigger than the than our parent block and um and also some checks about the the total uh terminal difficulty um that was needed for when we when we actually merged and um so we I think this is very gu specific but we don't start um synchronizing on new payload so if someone gives us a a block and we don't know anything about this block we don't know what the parent is um we need to ask the network for the parent um because it might be that we were offline for a bit so we didn't see the parent uh like arriving at us so we have to start a new new syn cycle but we kind of view this uh new payload uh new payload um method as kind kind of untrusted it's called with payloads that might or might not be wrong there might be uh uh uh payloads that um that are in the same slot um but one of them is the correct one one is the incorrect one they build on on on different blocks and so if we trigger the Sync here um so the synchronization is a very um very complex and uh inefficient process we need to ask the network for the parent um then the network tells us the parent and then maybe this parent is uh the parent of this block is something that we have already then we can say okay it's fine I'll just execute the parent and then I execute the new blog and then I'm I'm at the head and everything is fine but what if the parent of the parent is not there and so I have to ask the network again what's the parent of the parent and so on so triggering a sync is is a very complex method and um we only do this on on fog Choice updated the problem is uh on F Choice updated we don't have enough data to to trigger the syn well in theory we have in practice it's kind of murky but um what we do is when we get a new payload where we don't know the parent we will just put this in in storage and if the consensus layers is you need to you need to put this um this block as your head by calling the foras updated method um then we will go to the network and and start the syncing um so this is what we do here we we delay the payload import basically we just take this payload put it in our storage um wait for the next fure is updated if the next fure is updated references this block then we start the sync um and then so in order to execute all of the transactions correctly we need to have the state for for them right we need to have the um if if I want to verify that that a certain transaction spends the money correctly I need to know what was the balance of this account um before the transaction and uh so this is what we do here we check do we have the state for this parent block and if we have the state for the if we don't have the state for the parent block um we don't we we cannot do anything so we have to we have to just uh Warn and accept the accept the payload but um we cannot say whether this payload was valid or invalid because we don't have the uh the the state to execute it and um now comes the really interesting part um we we actually execute the block this is what we do in insert block without set head and um this method kind of only calls insert chain insert chain give gets a number of blocks and whether we want to set head set head means this is the new canonical block and I'm going to set my current header to this um and we don't do this in the in the um in new payload but only in for Choice uh in for Choice updated so um this insert chain is kind of a long method it's not really uh not really important what we do before we execute the block we have to verify the header um that it was uh that everything was um uh correctly built um I think we can go there so um we have the kind of the rules around the block are all written in the in the consensus package and here we have the different algorithms um that that we support so the beacon is for post merge click is for pre merge uh click is for like test networks uh and E is for premer and um so what we do here verify header um we just check is the terminal total difficulty reached and if it's not reached then we use um then we are pre merged and we have to use eash to verify that the header was was built correctly and otherwise we use the post uh method and here we have again some of the checks are duplicated here um that is because um um this part is uh used from all over the place um like the consensus um uh like the these consensus rules are the actual consensus rules um the other verification that we have is more like a did we unmarshal this block correctly or did something break there um this is where the actual consensus rules are verified and this is also verified during the sync it is not only like verified during during block uh during um uh during block execution so this is also whenever we get a block we have to verify it with this and um so if we go back uh back yes here we verify the headers and then we go through the blocks we we see if there's some reor it's not really really uh really interesting but the actual execution happens where's the ex execution sorry uh here um this is where we this is where we actually execute the block um basically we as I said in order to execute the block we need the state of the previous block all the balances all the smart contracts everything that is written in the smart contract uh we need this and this is uh this is what we do here we um State DB is kind of an abstraction um that we do on top of uh this state where we can we can write to it uh but it doesn't have to be flushed to diss so this is um basically we execute everything in memory and then uh we can call some function and then the the the the root of the of the mer p PR try is computed but we don't have to do this on everything uh like on every insert and um so this is where the actual block uh processing takes place in the block processor um and this is just the interface so um this method uh lives in core SL State processor this is where a lot of these um very internal uh execution um happens and um so what we do here before we before we can execute the block we have to um apply the Dow hard for uh to our state if that's if if we are on the Block that uh that enabled the Dow hard for um otherwise we create a new uh evm environment and uh set the transaction context and then we call this um apply transaction and in apply transaction we do apply message and uh this is where the ACT actual uh evm execution happens um so this is kind of where we also ended up in the last talk I'm pretty sure none of you have ever listened to the first talk uh but it's fine um this is where the actual um evm execution happens and before we have to before we can do any execution we have to verify that the sender of the transaction has enough money to cover um cover the transaction cost and this is what we do in in pre-check um where we we uh verify that the the the nons is correctly of this trans transaction and we verify that the Ender is uh the sender is an EA so you cannot send uh money from from account [Music] um it's just yeah we we hope there will never be like a hash collision between a an EA and a and a and an account but um now it's written in protocol and um then we have to verify that the base fee is correctly and um we have to also now with 4844 so as you can see here um it says blob hashes that means there's the also the logic for 4844 already implemented um already in master um uh but yeah so in addition to um to uh to the normal fees we also have to compute the the uh the blop fees and make sure that the the person um the sender has enough money to cover everything and so this is where we actually compute how much um it would cost and here is where we verify um that the sender has enough money so we get the balance from the sender and we uh we verify if they if they have enough money in the account and also we ver we verify if we have enough uh gas in the block left um because a the block can only be uh 30 million gas and uh so we need to make sure that if we charge this transaction we are not going to go over these 30 million um so this is what we also do here and then so this was the pre-check and um now we compute the intrinsic gas and then we prepare the state so we create the exess list we we make sure to put uh all the pre- compiles uh into into this list um so when someone calls a pre-compile it's cheaper uh than if they call like a normal address and then we have either a transaction that creates a contract or we have a transaction that calls something thing and um so if we have a contract creation we call the create op op code um and if we have a uh a normal call we just call the call up code and this is where we have to call um and um yeah so I think this is kind of where I'm going to stop because this is also covered by this other talk um but uh basically all of the functionality uh what happens within the evm is is in core slvm um this this also has the op codes uh uh for example this is the static call op code here or we have um the create two op code and um yeah so if you if you ever find yourself in a spot where you want to add an OP code um you need to do it here um yeah that's basically my my talk um I I still have some slides this is what I'm going to talk uh about in another another um Deep dive into go ethereum so there's hopefully going to be a a third one of these um where we talk about uh some of the other things thank you very much for listening we have room for a few questions if someone has a question from marus in the back yeah too hello thank you for your presentation um I was wondering from the Pro Prospect of layer twos um sometimes execution layer actually needs information from consensus layer is it actually possible for the execution layer to get information from the consensus layer an example of this would be what is the last EPO or block that has been finalized yeah um so this is uh the last finalized one that is and no there's no way for the consens for the execution layer to contact the consensus layer every the the data Flows In One Direction the execution layer tells the consensus layer what to do uh the ex uh consensus layer tells the execution layer what to do and um and uh checks for what what they said uh what you just talked about the finalized um hash in forch updated uh we get this update uh from the consensus layer and this sets the head block hash the save block hash and the finalized block hash um so uh you can kind of emulate this this uh two-way street in in in in that you just like provide all the data from the one side that the other side needs but this is not uh available for use by for example smart contractors right yes yes it is so we take we take we take this um and uh we Mark the blocks that are finalized and so if you go into your um uh if you have a dep and you say I want the U you can say you want want the latest block right uh get block by number latest or get block by hash latest or something and you can do the same thing with finalized you just say get blocked by uh hash and then finalized and you will get the final block this is um I think most of the clients already support this yes so G is in development for like eight or seven years now could what's the hardest part of maintaining it these days um the the biggest challenge is the state um because it's uh so big I think it's around 70 GB so this uh the state is um all of the accounts all of the balances the the basically the state try and right now we maintain um two copies of it we maintain one copy for basically only for reading um this is called a snapshot and we maintain Ain one copy um that is for computing the state route and that is the the state TR um and uh one thing that we've been working on for um for 3 years now is how to store the state try and previously we we stored every try note by hash um and now we we just um merged the path-based storage and uh this get rid of this problem where you uh you you you start your Noe and you sync and you have like an 800 7 750 GB note or something and then you let let it run for a while and it exponentially grows and then you can shut it down you can run the prune procedure delete a bunch of stuff and um then it's 750 GB again so um the problem there is we previously we stored data we stored some data that we did didn't need to store but we didn't know that this data is whether we can get rid of this data um now with the Pass based uh storage scheme uh we know exactly which data we can delete and um this will also like greatly reduce uh the size of our archive node um there are some preliminary numbers but it's an an order of magnitude less than what we need uh for storage right now and it won't grow as as quickly as um as with a with a hash based storage so yeah storring the state is the is the is our biggest issue because it changes so much and we have to recompute the try every time okay thank you let's give an Applause people thank
