New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Indexing Entire 2.4 Billion Transactions on Ethereum in 10 Hours | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

This talk covers learnings from building a general-purpose indexer which index every single transaction since genesis. There is also technical decisions when we have to deal with 7 billions records of data and how to process all of those data in less than half a day. Additionally, we will discuss the difference between batch data processing and real-time data processing, sharing best practices and strategies for both approaches. Speaker(s): Panjamapong "PanJ" Sermsawatsri Skill level: Intermediate Track: Developer Experience Keywords: Architecture, Scalability, Event monitoring, data, processor Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

okay so for our next talk we have penj to give us a talk on indexing entire 2.4 billion transaction on ethereum in 10 hours Let's welcome Pang J on the stage thank [Music] you oh okay thank you hello everyone uh my name is pan so I'm Thai so as a Thai uh welcome to Bangkok and Devcon here in Thailand all right so uh today I'm thank you okay uh so today I'm going to talk about uh indexing entire uh 2.4 billion transaction on ethereum in 10 hours okay um okay right now it may be like uh 2.7 billion right now but at the time it was a 2.4 okay let's get started so when talking about uh indexing on blockchain uh this is the con conventional way to do it so you have a database that tracks uh the latest block that you have processed and then you uh uh you you call the blockchain node and then get the data and then update the that block and so on So You Loop this uh uh until you get the latest block right this is a conventional way to to work on the blockchain indexer but uh looking on the product I'm working on is called Alat tress uh this one uh we we have some special requirements uh we want to index like uh all address in on the ethereum since Genesis and also we we want to get all the erc20 token events and then get those prices to get the what uh is called p&l the profit and loss that's uh that's what we want but how we do it we cannot do it in in a conventional way because it would take forever right uh to to read uh those uh 2.

4 billion transactions and then uh loop uh one by one so we need to uh think come up with the way uh to do the indexer in large scale so let's get back to the basic of the blockchain indexer actually it's just uh data processing right we've got uh input we've got uh then we put it into the process and then we've got output yeah so let's look into each component what we need from the input to be uh able to process it uh in at Large Scale is that we want it to be Compact and structured so this is the solution because uh pet give us like uh the compression uh also we have type in pcket as well c c pet format is like um CSV but better CSV and we found this project this is a very good open source project uh it read data on the ethereum Node and then write it into pocket file so that uh when we want to process we just uh read uh the file in in our file system that's is very lightweight we we don't need to uh call the node every time to get the data okay let's talk about the process how to do how to do the process in large scale so we have to do it in uh in a parallel way right there are couples of uh solution one is the Apache spark another is Apache beam uh we choose Apache beam because there is a Min service on the cloud platform and then output when we have process the data we have to write it somewhere right so uh that that uh destination that I have to write uh need to be able to scale horizontally me meaning that uh if I want to read a lot of and write a lot of data I can just uh create new instances and then uh I can read and write uh more there are a couple of solution to do the uh distributed dat database uh as you might guess that I selected the Google big table because I don't have to worry anything about infrastructure of the database all right now piece everything together and do some coding this this uh does not take uh 10 hours in i i in the coding uh it took me like uh two three weeks uh to work on all the components and understand how Apache beam works and actually write it uh this is uh what the final pipeline looks like so data comes from the top and then uh trickle down and then uh get the result and we wait around 10 hours and it finish just like that yeah um so we I I I uh ran it and it took like around 10 hours to index all the ethereum transaction all the uh erc20 token transfers and there's something more on it because we only index uh the current data right the big bat very very big batch of data and how to make it real time so uh I need to do it in a conventional way uh to to uh update uh to get the new uh blocks and then because I I don't need it to do at last scale anymore right I can just do it a traditional way to make it real time these are the numbers um it cost me about uh $350 and the number of the row that is written on the database in 10 hours uh was the uh 7.1 billion records and I scale the big table inance up to 20 incenses the good thing about big table is that when I finish uh writing the data I can scale it down so after I finish uh index those big data I scale out to just one to handle the normal we uh operation yeah that's that's the good thing about uh this whole thing yep I think uh that's all for me today thank you very much all right uh so we can ask some questions in the audience uh we can pass around the mic if there's any questions for p oh we do oh awesome we do have some questions on screen so let's go through the first one for eth node how did you host that to handle large read all right um for the in this process is the uh is the uh cryo right cryo reads from the node not not our not my indexer right so uh we use Min service and that is a dedicated uh server to to all the data yeah all right I guess we have one minute to take one more question uh where are you reading the archive data from how does reading from the node scale for you okay fortunate fortunately for my for my case I don't need the archive data I just need the transaction data the locks yeah from from the ethereum node that is just full node is suffice all right um so I think that concludes our talk for today

Automatic transcript — names and jargon may be misspelled.