# Scalable and sovereign EVM data: modern data engineering best practices

- Speakers: [Storm Slivkoff](https://streameth.org/speakers/storm-slivkoff)
- Channel: [Devcon 7 SEA](https://streameth.org/devcon_7_sea)
- Date: 2024-11-19
- Duration: 20:43
- Watch: https://streameth.org/watch/673cc8b7982f234a1257b1be

## Description

Collecting and analyzing large historical EVM datasets can pose a significant challenge. This has led many teams and individuals to outsource their data infrastructure to commercial 3rd-party platforms. However, over the past year a new style of data workflow has emerged, using entirely open source software and local-first processing. This new ecosystem of tools allow anyone to cheaply, easily, and robustly collect and analyze any EVM dataset from the comfort of their own laptop.

## Transcript

Okay, here we go. Hello everyone. Today I want to talk about sovereign data. And I want to start with a question. So let's say that you need to analyze Ethereum's history. What resources do you need to perform this analysis? And I don't just mean in a simple way, let's say that you need to analyze every block or every transaction or every state diff. What resources do you need to analyze these things? I'm here to tell you that you don't need very much. You don't need the cloud or any closed source software. You don't need a team of people to maintain your infrastructure. And you don't even need a database. The main takeaway from this talk is that you can analyze Ethereum's complete history locally on your laptop. And this is enabled by a lot of recent advances in open source data tooling that make this process really easy and really fast. And sometimes I describe it as it's like having BigQuery on your laptop, but it actually goes beyond that in a lot of ways because for many workloads this is actually faster and easier than using BigQuery. This is all enabled by an approach that I call data sovereignty. So the plan for this talk is I'm going to start off by defining what is data sovereignty, and then I'll explain why this is desirable, and then finally I'll explain how data sovereignty actually works and show you some live demos. The overall mission here is to convince you that data sovereignty is easy and fast and powerful. So in a nutshell, data sovereignty is just having full control over the data pipeline. And this goes further than just having an open data set. Really special things happen when the entire data pipeline is open source, it's modular using open standards, and it's available for you to run locally on your own machine. Even if you don't actually want to run it on your own machine, you can still get many benefits just by being adjacent to the tools and the ecosystems where these principles are prioritized. So what are the actual benefits? Just like with the terms open source software or decentralization, there's kind of two ways to answer that question. There's the ideological side, and there's the practical side. Ideologically, data sovereignty is the purest form of making data free and open. And your philosophy might be that you want complete control over your data. Data sovereignty is how you do that. On the other hand, there's lots of practical benefits. This type of workflow is often the easiest and fastest way to get the answers that you're looking for. And this is because it's powered by a rich open source ecosystem that is continually evolving. These days, there are hundreds of different tools you can use to analyze crypto data. And most of them weren't built for crypto. But we can use them thanks to the power of open standards and modularity. Another benefit of data sovereignty is that it can give you an extremely simple infrastructure with low maintenance burden. And this is relevant because most crypto companies don't have large data teams. They just have a single data person. And this means that we need systems that are easy for a solo operator to use. And as a final motivating point, I think that more data sovereignty will improve the EIP process. Ethereum is a really complex system. So when somebody proposes a change, it's important for us to evaluate that change using the relevant data. So what are all the downstream effects of this proposal? Is this proposal actually addressing an important problem? Sometimes these questions can only be answered using data. And in the past it's been kind of difficult to use Ethereum data. But my hope is that as data tools become better and better, this will lead to the EIP process becoming more data driven, and that will lead us to better EIPs. So how do we actually do data sovereignty? We basically just take a lot of the tools and best practices in data engineering and we apply them to the crypto world. And people have been doing this more and more the past couple years. So here's a flow chart of how I do most of my work. You start with an EVM archive node and then you use an ETL tool like Cryo to extract data sets from the node. That data gets saved in files on your computer, often in a modern format like Parquet. And then finally, you query those files using an engine like Polars or DuckDB. And that gives you results that you can use in your EIPs or your dashboards or whatever else. And one of the really nice things about this workflow is that it's very, very modular. You can run your own archive node, or you can use a third-party RPC provider. You can store these Parquet files on your laptop, or you could store them in S3. And you could query these files from your laptop, or you can run some sort of cloud computing engine. And this high level of flexibility is a direct consequence of having a modular ecosystem built on top of open standards. So let's zoom into the first step in this pipeline, which is data extraction. Last year I built this tool called Cryo for collecting blockchain datasets. Cryo can take any type of information available over RPC and turn it into a nice simple local data set. This can be simple stuff like blocks or transactions or more obscure things like opcode traces, JavaScript traces, really anything that's an RPC method. So a lot of this data can be really nested and messy when it comes out raw out of the RPC endpoint, but Cryo puts it into simple flat tables that are easy to consume. And you can use it as a CLI tool or as a Python library with the syntax shown here. So let me show you a demo of what cryo actually looks like. So let's see if this works. OK. Cool. Okay, cool. So the most basic usage is just collecting a vanilla data set. So let's say we want to collect all the logs over some block range. We can do cryo logs and then a block range, like let's collect from block 10 million, and then 100,000 blocks after that. And if we do that, it will collect this data set and save the output to a bunch of parquet files that are now on my laptop, which you can see here. And all of this is happening only on my laptop. I'm running a base chain node on the laptop as well, using Rath. And this interface is really simple. I could change to be a different data set, like blocks. So instead of logs, I do cryo blocks, and it would collect blocks over the same range. And you can see that it collects 100,000 blocks in about two seconds. So if you wanna list out the different data sets that Cryo can collect, we can do Cryo help data sets. List out a bunch of things there. And another thing to note is that Cryo is totally item potent. So let's say we want to collect a longer job, like collecting a million blocks instead of 100,000. I can just kill the job in the middle, and it doesn't matter. If I restart the job, it'll just pick up where it left off, and there's no corrupted files. There's no things to worry about in that regard. And so it's done I'll kill it again restart it and then finally when the job is complete you can see all the collected data and if I try to run it again cryo knows all the data has been collected and it just exits immediately and there's a lot of other options as well. And just for the purpose of time, I will sort of return to the presentation now. So great. So a lot of the magic of modern data tools is just having really good file formats. Parquet is the most common format used by modern data tools outside of crypto, and it's the default output format of cryo as well. Parquet can be seen as an extreme evolution of CSV. Most of you have probably seen CSV files before. They store tables. Every line in the file is a row in the table. Parquet stores tables as well, except it does so in a much more sophisticated way. It breaks the rows and the columns into separate chunks, and it compresses them, and it indexes them. So you get a very compact representation that can still be used for efficient queries that only read the subsets of the file that are relevant to you. So this leads to huge benefits in both speed and storage size. And there's now a huge ecosystem of tools that can create and read and update Parquet files. Here are some examples of common data sets that you might collect using Cryo and the size of those data sets when stored as Parquet. For example, it takes about 100 gigabytes to store the ERC-20 transfers from throughout Ethereum's complete history. And you'll need different data sets depending on what you're investigating. If you don't have enough room on your laptop, you can also grab a really cheap multi-terabyte thumb drive on Amazon these days. DATA SETS DEPENDING ON WHAT YOU'RE INVESTIGATING. IF YOU DON'T HAVE ENOUGH ROOM ON YOUR LAPTOP, YOU CAN ALSO GRAB A REALLY CHEAP MULTITERABYTE THUMB DRIVE ON AMAZON THESE DAYS. SO NOW I WANT TO SHOW A FEW DEMOS OF ACTUALLY USING THESE DATA SETS. AND I've collected in this directory blocks, ERC20 transfers, contracts, logs, and in each of these folders it's just a directory full of files. So if I do ls logs, it's just a bunch of files. And just to start off really simple, one thing we might want to do is gather all of the contract addresses. Maybe we want to label all of the contract addresses on Ethereum, or maybe we want to put them in a filter somewhere. So I have this script that just queries all the addresses from the contracts data set. So if I go bat demo zero, it's the main chunk in the middle that matters here. This contract equals thing. And this is using the Polar's library to scan a glob of files And it also selects which columns that you want to put into a data frame right here. It's just one column contract address So if we run this script Using that We need to go slash in front. Yeah, so it'll load all 67 million addresses from Ethereum in about one second. So if we time it, it's very fast, 0.77 seconds. And that's loading all of the addresses into memory. So now let's say we want to run a more useful query. Like, let's check what are the most common bytecodes for these 67 million addresses. Yeah, for these 67 million contracts. So I wrote another script. Demo 1. Actually, before I run it it let me show you it so similar story there's this contracts query used by Polar's library I'm selecting which columns I want here just code cache and then I'm adding this extra function value counts that's going to count the frequency of each bytecode. So if I run it, it'll load those contracts again, load their bytecodes, and here you can see it found almost 4 million unique bytecodes among these 67 million contracts, and it's also showing the frequency of each bytecode. Here are the top 10. And it was able to do this in about three seconds. Oh, yeah, the screen's a bit small, so it can't quite fit. And then another thing that we can do is really efficient processing of ERC-20 data. So let's say we want to look at WBTC, and we want to load all of the WBTC transfers. So here's another script, demo. Here it's still very simple. We're scanning a blob of files. Here I'm just reformatting the addresses as hex. But, like, let's run this script. And you can see very quickly it loads about 8 million WBTC transfers in about one second. And it also loads all of the relevant metadata about each transfer, like the transaction hash, block number, sender, receiver, all that. And also note that this allows us to process datasets that are much larger than our memory size. The ERC20 transfers data set is 97 gigabytes, and that's larger than the memory I have on this laptop. And finally, let's compute all of the WBTC balances of all holders. And so this final script is a bit more involved. It's basically finding... You're not going to be able to see much on this tiny little screen, but it's basically computing the inflows and outflows to each address and aggregating by address. So if we run that, we can see we basically compute all of the holders of WBTC, and here we're displaying the top 10. And it also computed for 100,000 additional holders, and it does all that in about a second and a half. And this is not specific to WBTC. You can do this for the majority of ERC-20 tokens. You can get really fast data on the distributions of those tokens. Okay, so let me return to the slides. Here we are. Okay, so those are just some contrived examples, but things can get a lot more complicated in practice. I wanna finish with a couple of final thoughts. contrived examples but things can get a lot more complicated in practice. I want to finish with a couple final thoughts. First thing is data sovereignty is not an all-or-nothing thing. You can think of it as a spectrum and just like there are different levels to decentralization, there are different levels to sovereignty. Maybe something is closed source but but it uses open standards. Or maybe it's open source, but you can't run it locally. And in the same way that you don't always need 100% decentralization, you don't always need 100% sovereignty. I use plenty of tools that are not sovereign all the time. You should always be using the right tool for the job. And the final thing I want to say is that sovereignty is a uniquely Ethereum phenomenon. Ethereum's design philosophy prioritizes introspection of what's happening on-chain. And this has enabled a rich ecosystem of tools that collectively make the sovereign workflow possible. This is the culmination of many years of work by many idealistic people, so I think we should all appreciate that we're in a pretty awesome situation with respect to data. And I think it's only going to get faster and easier and more capable as time goes on. So that's everything I wanted to share. Thanks for listening. Thank you very much, Storm, for this insightful talk. I guess everyone who's new to data or who's already familiar learned quite a lot today. And also you killed it with the live demo. The live demo devils were away from us today. Yeah, it was a little nerve-wracking. Right. So I'll move on with the questions. The first one from the leaderboard is, is it suitable as a subgraph replacement consumed by end-user frontends? So the workflow that I'm showing here is kind of more suited to historical data analysis rather than live data analysis. There are some extensions that could happen in the future, and it's something we've thought about, like for Cryo and for other related tools, but the focus here is really on historical data analysis. The next one is if you can use Cryo to serve my app which needs near real-time data. That's basically the same question. So same answer. Next up, what does it take to build similar tooling for L1 and L2? The tooling is exactly the same. So this was actually technically an L2 demo because I'm running a base chain node on this laptop and I was querying all the data there. JSON RPC is kind of like a common lingua franca of VM nodes. So anything that speaks JSON RPC, you can do the same exact sort of flow with cryo. And then Polar's is even more general. It can process any Parquet file you want. How hard is it to define a new dataset? It's actually not that hard. So in cryo, for example, each dataset is just a single simple file. It's sort of like defining the flow of like take the RPC data, manipulate it a little bit, and then like output it as a row. Cryo already has a ton of different data sets though, so you probably don't need to define new data sets. What is the most common contract on ETH? I'm not sure. I think it's probably either a proxy or like a Uniswap pool. I think I can answer that. In Sourcefy, we have a data set as well, and I think it's Gnosis Safe proxy. Ah, nice. That makes sense. At least to our data set. Why Solana is not data sovereign? Well, it's kind of the reasons I listed. So Solana, it just hasn't really been a priority to make introspection of what's happening on chain really visible. The scale is also an issue with Solana. Solana data is basically too big to fit on any local machine. So you're not going to be able to do local processing there. All right. That's it from our questions. Let's give a big round of applause.
