Using Reth Execution Extensions for next generation indexing by Alexey Shekhirin | Devcon SEA
Devcon·Tue, Oct 7, 2025, 12:00 AM
Recently, Reth and Geth released the ExEx and live tracer features, respectively, which share similar functionalities. Both provide real-time, detailed access to chain and state events. As ExEx developers begin to persist this data and explore ways to make it accessible to users, new questions arise: how can we best serve this data to users, and what might the indexers of the future look like? Speaker(s): Alexey Shekhirin Skill level: Intermediate Track: Core Protocol Keywords: Layer 1, Developer Infrastructure, Tooling, plugin Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/
Transcript
[Music] hey everyone I'm Alexi I work uh on W at ttha and today I'll be talking about W execution extensions for Next Generation indexing so let's start with what is indexing why do we need it uh ethereum is a database in itself it has uh history State and receipts and it sounds like it's sufficient for all our needs but it's not because history is too large and um as an example when you want to query all the transactions for a user you need to go through all the transactions in the history same with State uh you can query the account balance and account nons very easily but for example to query the erc20 balance you need to actually index every smart contract for the erc20 and uh calculate how many tokens does the user have um receipts and logs are good they work but unfortunately sometimes they lack all um the required data in them and uh you still need to know what receipts to query so you can have some transactions that meet the locks about the deposit for example into the Bon chain uh deposit contract but you need to query uh the transactions for uh this for this contract so people usually build indexers uh indexers are third party solutions that uh store the data Elsewhere for example it's ether scan that we all use uh it's Dora um and also wallets wallets are indexers they usually Outsource it to someone else like infura ether scan whatever API but to show the token balances you also need it and uh if you think about it um a rollup can also be a kind of an indexer because you need to keep track of the deposits from L1 to L2 you need to uh derive the data from L1 to get the L2 State uh for example when syncing um so yeah it's an outof note uh infrastructure that indexes onchain data uh to get some useful data of chain um so how do indexers work today you as an indexer developer have several choices and the most obvious One is using the Json RPC Json RPC has has the eth Subscribe method that allows you to get notifications for new blogs uh that arrive this will get you the transactions and you can also Trace these transactions later you can inject your code into the node source code for example you can have G or W or nethermind code you Fork it you inject your code for storing the uh data from the from the transaction into your storage of choice pogress cka whatever and uh you process it later and you can build on top of the existing node components and create your own indexing node that runs in the same uh binary in the same process as the node itself so let's go through all of them and uh talk about pros and cons uh Json RPC subscriptions very easy every node has it this is the uh the thing that's almost speced so people uh expect the behavior from every note to be the same um the cons of this approach is the poor performance mostly due to Json RPC uh serialization and data transfer because you need to serialize to Json you need to uh make an HTP requests then you need to deserialize all the things um it's difficult to handle the chain reorgs because when your chain reor you get a notification on the eth subscribed um e subscribed web circuit connection but this notification doesn't tell you a lot about what was reverted and what was committed now you need to figure it out by your own and also it requires the ad hoc solution some like rust python JavaScript uh code that trans alone sign your note uh listens to this Json RPC method and uh stores this data somewhere okay we have like a better idea uh we can inject our code into the node Source we solve the problem with overhead on ization and the data transfer you basically do not leave the memory that the note uses and um all data is potentially accessible so you have the access to the database um to the transaction pool to everything that the node has access to and um G has it today it's called life tracing and um as you can tell it's very easy you add your own uh Dogo file to some directory um you run gu as normally and it produces your locks the cons of this approach um maintaining the fork of an original uh repository so you need to for G you need to add um a dogo file into some directory but if for example G introduces some new feature you need to um pull the changes you need to uh resolve the conflicts and then you can uh start running your note again and the con is also that you aside to the programming language that you use uh that the note is written in so in case of gu is go so we have our own answer to that and it's execution extensions so let's talk about what day um executions extensions require um allow you to consume a reor aware stream of notifications coming from the node basically you get access to all node components same as in the approach with gu you get access to evm database transaction pool payload Builder you use share memory for communication so the node sends a notification about chain commit it has one block this block is in memory and all xaxis for example you can have five execution extensions they share the same memory so there is no copying uh going on and it's also compiled into the same binary as W and um it's it's run in the same process so as I said there is no forking uh I didn't say this there is no forking because we use W as a library um that's the main I think difference that we try to um that we try to talk about more that um WTH is an SDK not just a note and you can import WTH you can use it in your own small project uh 100 lines of code uh 1,000 lines of code doesn't matter and uh when uh W updates you just do cargo update and it updates your dependencies you also have the access to all the node components as I said and uh what's also very cool that you get access to the state divs uh right from the evm execution so the evm finished executing the block it um it produce some State diffs and TR divs and you can consume them the issue is that you are try you ATT tie to the programming language um of Wrath which is rust which is not bad uh so as I said you can do whatever uh you want as like an indexer you can do an indexer a rollup for example example a base rollup an AVS or an Oracle and uh you use the power of the node Builder uh in SDC to introduce a new RPC method for this uh indexer that serves your data to extend the payload Builder or add a new CLI argument um to somehow configure your indexer without modifying the source code this is an example of a of an XX um what happens here is that you consume the stream of notifications there is three variants for the enam chain commited chain reor and chain reverted and and as you can see in the reor we have old and new chain so you can uh undo all the all the state diffs from the old chain and apply all the state diffs from the new chain and then the uh node is started in the FN main entry point function which is like 10 lines of code and and it runs the full R node just as normal R node but with your extension uh that node send data to so people today are are already building stuff with xais we have uh Tao experimenting with the uh base trup built on x-axis we have Kona which is an opls uh project for deration of uh L2 blocks from the L1 and it's also an nexx uh we have Shadow uh working on Shadow logs when you want to have additional logs in your app uh that are not in the source code of the uh contract itself and you on every uh new transaction with this contract you modify the VM execution of it uh and um V we am also doing some stuff and we and they have like three four blocks uh four block posts in their uh block about it and rups are also can be seen as just execution extensions as you can see and uh that's a future I'm personally very uh looking forward to that people built into one piece of software where many plugable pieces and you run a rollup inside your own note without communication via some API so let's build an indexer uh what we want to do we want to index uh the Bon chain deposits for for example and serve them as a list of top depositors to the contract through an API uh deposits are recorded in the events of the contract we can decode them and the API will be the node Json RPC and we extend it uh using our own method that we will build and this is the screenshot from theer Scan they already have this uh so the idea is that the power is in the stack that we built and uh we can use uh as I said X access for listening to new blocks uh we decode this uh these transactions and chain events um using using alloy for storage I will use sqa just because it's the safest uh choice for a small project and by database uh for the API I will use a custom Json RPC method um and we will bundle it all together with the node Builder and test using cast from Foundry to send the RPC method to send the RPC uh request to our node so listening to new blocks basically the same as I showed before same bowler plate you have the notifications Channel you consume it and you react on different events uh the the chain events we have two uh definitions here we Define the deposit contract address which is uh done with the address micro from alloy uh which represents the address in memory efficient way and not as a string and uh here the coolest thing is the soul macro which passes the uh solidity source code and generates the RAS code from it basically you'll have a stru deposit events with this Fields PP key withdrawal uh credentials Etc and it will not generate the code into a separate file but we'll have it in the um same file accessible through your LSP or whatever uh you use and during the compilation it will uh generate this code uh for storage we use SQ light and uh what we store is uh one row per uh per deposit so every time we have a reor we have a revered chain and we delete all transactions that that in the reverted chain from the database and we do it by the hash for the commit we do a bit more we go through uh blocks in the chain because for example if it's a 2d preor block you will have two blocks in the reverted and two blocks in the committed chain you go through the blocks you go through transactions with receipts you go through the logs in each transaction and you also filter only those transactions that are uh sent to the beon chain deposit uh contract you filter the locks uh you decode the locks and uh then you insert this data with the amount uh from the loog into your database API we Define a trade uh that has a name space of Devcon and the method of depositors um there is one argument count for limiting the output length uh how many top depositors do we want to uh get and the all the the this method does is sare in the database uh with some simple SQL web to uh we have node Builder to bundle it all together we create a new ethereum node uh we install the XX that we build we extend the RPC with our module that we buil and we start the Lo the node so when the nodee is up and running it needs to sync but if you already had the sync re node you can use it because it's still W mainnet ethereum node you just have an extension on top of it so if you already had a sync node it will not uh have anything in the database but it will backfill uh all the required blocks uh from Genesis and then you will have the sorry and then you will have the depositors that you can query with cast and cast RPC is just a common to query the Json RPC so the result is one file with less than 200 lines of code for building the indexer and uh the power here is that you have the whole W code base with thousands of lines of code that's uh hidden by the UT and um in the corner there is a QR code you can scan it and check out the repo so what else can you build with xaxis uh we have a couple ideas that we would like people to experiment with and uh if you have any other ideas we are very welcome and um we would love to hear it from you so the remote XX is about streaming the data out of the node into some data Lake for example you have a a JPC server or you have some Kafka or any other queue and you just simply stream this data out and consume it on the other end I believe this is what uh companies like ether scam will do because they will not plug the uh the indexing Parts into the XX because they have a more complicated system Oracle XX uh you listen as an XX to some offchain data source for example you query the New York Times uh web site you attest to this data as one user uh of this XX uh so you sign uh that you saw this kind of data uh you gossip this stations using the dq5 uh topic to other people running the same exx and what you have you have a small network of uh of oracles trying to agree on something and when they come to consensus uh when the Quorum is reached they post their data on chain with these uh signatures so you can tell that uh it's actually signed but by by all these people who uh said that they oracles and it's very similar to how avss work and you can build avss with with xais based trups um you expose your own RPC to send transaction you do not poess transactions to L1 but instead you execute them on top of the state of your rollup then you bdch them and you submit the the transactions and the state to L1 and uh on the QR code we also have a bunch of examples that we build for people to play with uh they're not production ready but they uh serve as as an inspiration for people to start working with xaxis that's it for today uh thank you and let's build thank you very much for the wonderful talk Alexa so let's uh move to questions as a reminder please scan the QR code and submit your questions so this is very Dev friendly but uh you still need to run a full W node right even just for running simple indexer yeah that's true that's basically true what yeah um you can either run an archive node to get access to all history and all change sets or you can have a full rest node yes okay how does backfill work right so when you start your XX there is an option to subscribe to Notifications telling what's the latest head that you saw is and for example if you start your XX on block 21 million but you didn't have any data in your XX database you passed a block zero and uh it started sending first the blocks and um and change sets and state divs from block zero to 21 million and then it sends all the data from block 21 million to the tip uh what's the delay before we receive an XX notification ASAP ASAP um um basically you execute and once it's executed you send a notification so we send the notification and you receive it um matter of milliseconds yeah okay uh did you consider building this in some language agnostic but still in process way that's a good question um we believe that it should be built by the community and it can be a SAS uh separate project whatever basically the idea is to execute some logic in your XX that's written in for example uh go that that's compiled to to wasm and uh you send this notification to the wasm uh runtime it can be done and we have a PR on uh example rep to do it with bom and it's working but it's not merged because it's not finished okay will the XX run is a separate process or would it just add some custom logic into W and run only one W process all right so default way to run x-axis is in the same process what you get from this is a separate uh sorry shared memory you do not waste time on somehow communicating between this separate process and re but as I said you can run a a remote XX that allows you to just stream the data out of your Noe and um you consume it with your code of choice python JavaScript whatever if in something similar to its subscribe way but you can choose what it should be it can be protu it can be some uh Json whatever uh what do you mean by reorg aware notifications yeah so when you have a note running sometimes it reworks meaning that the blog that's canonical now um will be replaced by another block or last two three for whatever blocks can be replaced by different blocks so your exx to maintain the canonical State needs to be aware of these uh rears and these changes so when it happens you receive a notification saying hey this is the all chain and the all chain is for example these two blocks that we uh deleted from the database and these two blocks are the U that we committed to the database and you can apply this to your separate a database that you maintain for example with your state of uh deposit uh deposits to the contract okay uh how much time does it take to backfill XX for simple use cases such as logs and for more advanced use cases like internal contract State yeah so it doesn't really matter I think what you do here because you receive the state divs and the logs anyway what matters is the evm execution time because what happens when you backfill you run pure evm execution from block for example 0 to 21 million and uh currently I'm not sure what's the speed of this the peak speed I would say it's um some number of gigas mentioned uh per second and uh this is the speed that you can get uh new updates at how do you handle reorg it's handled inide R nodes and we basically maintain a structure the tree that says this is the canonical chain this is the side chain and at some point the side chain becomes canonical so we uh revert the canonical chain and make the side chain canonical uh that's a rough explanation and what you get as the nexx developer is a nice de old new chain uh does back filling require running XX in archive mode or is a full Noe enough it depends on how deep is the back fill so in uh in the full note we still store some change sets and uh if it's for example 20 blocks deep it's fine but if it's from Genesis you need an archive notes how do offchain oracles coordinate data sources that's up to you and um it should be some protocol that agrees uh what to query uh how to agree uh how to communicate this data between each other Etc um if uh did you see that g uh has a PR for this now too yes I and unfortunately it's closed now because Peter wants to move it to his own small repo I respect this um this is good yeah I like it okay um if XX runs as an independent process what's the difference um between that and other indexing tools such as the graph or other ETL tools you basically can choose what methods of communication it uses and for example you can use protuff that's more memory efficient and space efficient uh what are invariance of the XX interface are notifications guaranteed to be ordered and contiguous yes okay um similar to the plugin in hyper oh okay that seems to have been uh does backfill still work if you enable pruning again depends on the depth of the backfill okay um can you alter execution with XX for example execute some custom Logic for each store uh yes you can you can override it with a node Builder using the custom evm executor uh it's independent from the XX and the XX will get the result of your evm execution and uh I think a final question uh how is XX compared to hook from G or Plugin from nethermind right G hooks I think require forking and plugin from nether Minds is a cool idea that works via Dynamic libraries and I think it's very similar in terms of performance and just a different approach all all right well I think that's it for questions thank you very much LSA [Applause]
Automatic transcript — names and jargon may be misspelled.