# Scalable and sovereign EVM data: modern data engineering best practices | Devcon SEA

- Channel: [Devcon](https://streameth.org/devcon)
- Date: 2025-10-07
- Duration: 20:43
- Watch: https://streameth.org/watch/yt-bKrnOnfx9io
- YouTube: https://www.youtube.com/watch?v=bKrnOnfx9io

## Description

Collecting and analyzing large historical EVM datasets can pose a significant challenge. This has led many teams and individuals to outsource their data infrastructure to commercial 3rd-party platforms. However, over the past year a new style of data workflow has emerged, using entirely open source software and local-first processing. This new ecosystem of tools allow anyone to cheaply, easily, and robustly collect and analyze any EVM dataset from the comfort of their own laptop.

Speaker(s): Storm Slivkoff
Skill level: Intermediate
Track: Developer Experience
Keywords: Developer Infrastructure, data, analysis, Developer, Infrastructure

Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon
Learn more about devcon: https://www.devcon.org/
Learn more about ethereum: https://ethereum.org/ 

Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more.

Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. 
Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024.
Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

## Transcript

[Music] okay here we go hello everyone uh today I want to talk about Sovereign data and I want to start with a question so let's say that you need to analyze ethereum's history uh what resources do you need to perform this analysis and I don't just mean in a simple way let's say that you need to analyze every block or every transaction or every state diff what resources do you need to analyze these things I'm here to tell you that um you don't need very much you don't need the cloud or any close Source software you don't need a team of people to maintain your infrastructure and you don't even need a database the main takeaway from this talk is that you can analyze ethereum complete history locally on your laptop and this is enabled by a lot of recent advances in open source data tooling that make this process really easy and really fast and sometimes I describe it as it's like having big query on your laptop but it actually goes beyond that in a lot of ways uh because for many workloads this is actually faster and easier than using big query this is all enabled by an approach that I call data sovereignty So the plan for this talk is I'm going to start off by defining what is data sovereignty and then I'll explain why this is desirable and then finally I'll explain how data sovereignty actually works and show you some live demos the overall mission here is to convince you that data sovereignty is easy and fast and Powerful so in a nutshell uh data sovereignty is just having full control over the data P line and this goes further than just having an open data set really special things happen when the entire data pipeline is open source it's modular using Open Standards and it's available for you to run locally on your own machine even if you don't actually want to run it on your own machine you can still get many benefits just by being adjacent to the tools and the ecosystems where these principles are prioritized so what are the actual benefits just like with the terms open-source software or decentralization there's kind of two ways to answer that question there's the ideological side and there's the Practical side ideologically data sovereignty is the purest form of making data free and open and your philosophy might be that you want complete control over your data data sovereignty is how you do that on the other hand there's lots of practical benefits uh this type of workflow is often the easiest and fastest way um to get the answers that you're looking for and this is because it's powered by a rich open- source ecosystem that is continually evolving uh these day these days there are hundreds of different tools you can use to analyze crypto data and most of them weren't built for crypto but we can use them thanks to the power of Open Standards and modularity another benefit of data sovereignty um is that it can give you an extremely simple infrastructure with low maintenance burden and this is relevant because most crypto companies don't have large data teams they just have a single data person and this means that we need systems that are easy for a Solo solo operator to use and as a final motivating point I think that more data sovereignty will improve the e IP process ethereum is a really complex system so when somebody proposes a change it's important for us to evaluate that change using the relevant data um so what are all the downstream effects of this proposal is this proposal actually addressing an important problem sometimes these questions can only be answered using data and in the past it's been kind of difficult uh to use ethereum data but my hope is that as data tools become better and better this will lead to the EIP process uh becoming more data driven and that will lead us to better eips so how do we actually do data sovereignty we basically just take a lot of the tools and best practices in data engineering and we apply them to the crypto world and people have been doing this more and more the past couple years so here's a flowchart of how I do most of my work you start with an e an evm archive node and then you use an ETL tool like cryo to extract data sets from the node that data gets saved in files on your computer um often in a modern format like parket and then finally you query those files using an engine like polers or duct DB and that gives you results that you can use in your eips or your dashboards or whatever else and one of the really nice things about this workflow is that it's very very modular you can run your own archive node or you can use a third party RPC provider you can store these parket files on your laptop or you could store them in S3 and you could uh query these files from your laptop or you can run some sort of cloud computing engine and this high level of flexibility um is a direct consequence of having a modular ecosystem built on top of Open Standards so let's zoom into the first step in this pipeline which is data extraction last year I built this tool called cryo for collecting blockchain data sets cryo can take any type of information available over RPC and turn it into a nice simple local data set this can be simple stuff like blocks or transactions or more obscure things like op code traces JavaScript traces uh really anything that's an RPC method so a lot of this data can be really nested and messy when it comes out raw out of the RPC endpoint but cryo puts it into simple flat tables that are easy to consume and you can use it as a CLI tool or as a python library with the syntax shown here so let me show you a demo of what uh cryo actually looks like so see if this works Okay cool so um the most basic usage is just collecting a vanilla data set so let's say we want to collect all the logs over some block range uh we can do cryo logs and then a block range like let's collect from block 10 million um and then 100,000 blocks after that and if we do that it will collect SCT this data set and save the output to a bunch of paret files that are now on my laptop um which you can see here and it's and all of this is happening only on my laptop I'm running a base chain node on the laptop as well using W and this interface is really simple I could change to be a different data set like blocks so instead of logs I do cryo blocks and it would collect blocks over the same range um and you can see that it collects 100,000 blocks in about two seconds um so if you want to list out the different uh data sets that cryo can collect we can do cryo help data sets um it'll list out a bunch of things there um and another thing to note is that cryo is totally item potent so let's say we want to collect a longer job like collecting a million blocks instead of 100,00 um I can just kill the job in the middle and it doesn't matter if I restart the job it'll just pick up where it left off um and there's no corrupted files there's no things to worry about in that regard and so it's done I'll kill it again restart it and then finally uh when the job is complete um you can see all the collected data and if I try to run it again CYO knows all the data has been collected and and it just exits immediately and there's a lot of other options as well um and just for the purpose of time I will sort of return to the presentation now so great so a lot of the magic of modern data tools is just having really good file formats par is the most common format used by modern data tools outside of um crypto and it's the out it's the default output format of cryo as well paret can be seen as an extreme evolution of CSV most of you have probably seen CSV files before they store tables every line in the file is a row in the table par stores tables as well um except it does so in a much more sophisticated way it breaks the rows and the columns into separate chunks and it compresses them and it indexes them so you get a very compact representation that can still be used for efficient queries that only read the subsets of the file that are relevant to you so this leads to huge benefits in both speed and storage size and there's now a huge ecosystem of tools that can create and read and update parket files here are some examples of common data sets that you might collect using cryo and the size of those data sets uh when stor stored is paret uh for example it takes about 100 GB to store the erc2 transfers from throughout ethereum's complete history and you'll need different data sets depending on what you're investigating if you don't have enough room on your laptop you can also grab like a really cheap multi-terabyte thumb drive um on Amazon these days so now I want to show a few demos of actually using these data sets um and I'm going to do this using entirely local process processing on my laptop so let's return to endline and okay so I've collected in this directory blocks erc20 transfers contracts logs and in each of these folders it's just a directory full of files so if I do LS logs it's just a bunch of files and just to start off really simple um one thing we might want to do is gather all of the contract addresses maybe we want to um label all of the contract addresses on ethereum or maybe we want to put them in a filter somewhere so I have this script that just queries all the addresses from the contracts data set so if I if I go bat demo zero um it's the main Chunk in the middle that matters here um this contract equals thing and this is using the polar's library to scan a glob of files um and it also selects which columns that you want to put into a data frame right here it's just one column contract address um so if we run this script using like that oh we need to go slash slash in front yeah so it'll load all 67 million addresses from ethereum um in about 1 second so if we time it it's very fast 77 seconds and that that's loading all the addresses into memory so now let's say we want to run a more useful query like let's check what are the most common bite codes for these SE for these 67 million addresses or for the yeah for these 67 million contracts so I wrote another script um demo one actually before I run it let me show you it so similar story um there's this contracts query used by uh polar's Library I'm selecting which columns I want here just code cache and then I'm adding this extra function value counts that's uh going to count the frequency of each bite code so if I run it it'll load those contracts again load their bite codes and here you can see it found almost 4 million unique bite codes among these 67 million contracts and it's also showing the frequency of each bite code here are the top 10 and it was able to do this in about 3 seconds oh yeah the screen's a bit small so yeah can't quite fit and then um another thing that we can do is really efficient processing of erc20 data so let's say we want to look at wbtc and we want to load all of the wbtc transfers so here's another script uh demo here it's still very simple we're scanning a blob of files here I'm just reformatting um the addresses as hex but like let's let's run this script and you can see very quickly it loads about 8 million wbtc transfers in about 1 second and it also loads all of the relevant meta metadata about each transfer like the transaction hash block number sender receiver all that and also note that this allows us to process data sets that are much larger than our memory size the the erc20 transfer data set is 97 GB and that's larger than the memory I have on this laptop and finally let's compute um all of the WT wbtc balances of all holders and so this final script is a bit more involved it's basically finding you're not going to be able to see much on this this tiny little screen but it's basically Computing the inflows and outflows to each address and aggregating by address so if we run that we can see we basically compute all the holders of wbtc and here we're displaying the top 10 and it it also computed for 100,000 additional holders and it does all that in about a second and a half and this isn't this is not specific to WB TC you can do this for the majority of erc20 tokens you can get really fast data on the distributions of those tokens okay so let me return to the slides here we are okay so those are just some contrived examples uh but things can get a lot more complicated in practice I want to finish with a couple final thoughts uh first thing is data sovereignty is not an All or Nothing thing you can think of it as a spectrum and just like there are different levels to decentralization there are different levels to sovereignty uh maybe something is closed Source but it uses Open Standards or maybe it's open source but you can't run it locally and in the same way that you don't always need 100% decentralization you don't always need 100% sovereignty I use plenty of tools that are not sovereign all the time uh you should always be using the right tool for the job and the final thing I want to say is that sovereignty is a uniquely ethereum phenomenon ethereum's design philosophy prioritizes introspection of what's happening on chain and this has enabled a rich ecosystem of tools that collectively make the Sovereign workflow possible this is the culmination of many years of work by many idealistic people um so I think we should all appreciate that we're in a pretty awesome situation with respect to data and I think it's only going to get faster and easier and more capable as time goes on so that's everything I wanted to share thanks for [Applause] listening thank you very much storm for this insightful talk I guess everyone who's new to data or who is already familiar learned quite a lot today and and also you killed it with the live demo the live demo Devils were away from us today yeah it was a little nerve-wracking right so I'll move on with the questions um the first one from the leaderboard is is it suitable as a subgraph replacement consumed by end user front ends um so the workflow that I'm showing here is kind of more suited to historical data analysis rather than live data analysis um there are some extensions um could happen in the future and it's something we've thought about like for cryo and for other related tools but the focus here is really on historical data analysis the next one is if the if you can use cryo to serve my app which needs near realtime data uh that's basically the same question so same answer right um next up what does it take to build similar tooling for L1 and L2 uh the tooling is exactly the same so this was actually technically an L2 demo uh because I'm I'm running a base chain node on this laptop and I was quering all the data there um R Json RPC is kind of like a common lingua franka of evm nodes so anything that speaks Json RPC you can do the same exact uh sort of flow with cryo and then polers is even more General it can process any parket file you want how hard is it to Define any new data set um it's it's actually not that hard so in in cryo for example each data set is just a single simple file um sort of like defining the flow of like take the RPC data manipulate it a b a little bit and then like output it as a row um cryo already has a ton of different data sets though so you probably don't need to Define new data sets uh what is the most common contract on eth uh I'm not I'm not sure I think it's it's probably either a proxy or like a Unis SWAT pool I think I can answer that uh in sourcify we have a data set as well and I think it's gnosis safe proxy ah nice that makes sense at least to our data sets why Solana is not data Sovereign well it's it's kind of the reasons I listed so Solana it it just hasn't really been a priority um to make introspection of what's happening on chain really visible um the scale is also an issue with salana um it it's basic the salana data is basically too big to fit on any local machine so you're not going to be able to do local processing there all right that's it from our questions let's give a big round of applause
