New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Blockchain Storage - The Ultimate Source of Truth for Analytics - Tomek Mierzwa

ETH WarsawTue, Oct 7, 2025, 12:00 AM

Storage diffs are the unsung heroes of the EVM compliant blockchains. They ensure all nodes agree on the current blockchain state, enabling smart contracts to function correctly, saving resources, and enhancing security by detecting unauthorised changes. Obtaining storage diffs is a complex task that demands considerable computational resources, making it less accessible to the average consumer. In the current state of web3 analytics, we are hamstrung as we're unable to reconstruct and thus understand exactly what is happening on chain without this critical component. At the same time, they are incredibly difficult to decode making the information they hold not available to the majority of the network users. If you want to know the value at a specific memory location, you can use web3's getStorageAt function. However, when it comes to tracking storage changes within a block, it's a challenge. You typically only see the values before and after the entire block, which is less than ideal. To make matters more difficult, many memory locations are calculated as hashes of slots and keys, making it difficult to iterate over hash-maps without knowing the key in advance. Token Flow has been working on making using storage easier. This presentation explains how to do analytics without having to rely on calls and events alone. 🧜🏻‍♀️ ETHWarsaw is a series of educational and entertaining events for an active community of blockchain builders, developers and enthusiasts with focus on Ethereum-related tech. Once a year, we organize a large conference and hackathon for the community in the center of the Polish capital with speakers from the best web3 projects and participants from all over the world. Follow ETHWarsaw on social media for the latest updates! X (Twitter): https://twitter.com/ETHWarsaw LinkedIn: https://www.linkedin.com/company/ethwarsaw Telegram chat: https://t.me/joinethwarsaw See you all at our events in Warsaw 🙌🏻

Transcript

talk about blockchain storage uh our platform is actually designed for very uh complex use cases uh use cases which cannot rely on events and calls which are like the typical source of data for blockchain analytics today so uh we provide the multi-chain data which is natively like multi-chain so our dat sets are from the very beginning like multichain so whenever you are interested in querying data across multiple Chains It's like natively supported in our uh data sets and we provide the same data as others like all the events calls transactions blogs full history of everything but we also focus on storage information storage which is like the the major topic of this stock I will uh focus on uh storage and we'll try to convince you how you can analyze data using storage in completely different way uh we provideed also some other types of data that can be used for uh various Edge case analytics like for example full history of reverts not only reverted transactions but also reverted events reverted changes in storage because many things can happen happen in blockchain execution that then got reverted when the whole transaction is reverted normally it's not preserved it's not stored in the note uh you can find it in our data we even provide the full history of Trent storage the thing that was uh uh introduced with the Denon uh hard Fork so storage that is temporal something that's uh stored only during the transaction execution and not it's not preserved up transaction ends uh still if you want to have like full insight into the history of execution you need this information and you will find this information in our data sets so uh as I said we provide multi-chain data which uh has all those things that you are very used to for blockchain analytics plus the storage information that I would like to focus a little bit more during this uh this short talk so as you know very well we in in blockchain we basically have is everything is about storage blockchain is a machine that stores some information in time in very secure way and this storage is decentralized is permissionless is something that people can use to verify many things like for example balance of your tokens and you have actions actions that change this storage actions read information from the storage and a right to storage so basically uh every action in blockchain history is about doing some changes in storage and every action or many actions today they emit events this is like the standard thing that was uh introduced at the very beginning of the concept of evm creation of concept of EDM that people will need to analyze data and because people will need to analyze data ethereum creators they introduce this concept of events with are kind of logs that developers can Implement in their contracts and they emit some information that then can be easily quered and used for analytics and it is like the basic very popular approach to analyzing everything in blockchain history we look at if events which are actually results of some actions because when you execute transaction you actually uh execute some contracts and in this contract code you have actions that emit some events so if I transfer token I emit event transfer event that is stored and easily accessible I can then get history of all events this is like very standard way of analyzing uh history something that we are very very used to but there are many problems with events they uh are becoming more and more hard to use first because events are costly like emitting events it costs you money so more and more contracts they try to uh like minimize the gas spend for the contract execution and they try not to emit event some contracts today they some some protocols they completely like stop emitting any events because it's not obligatory like some standards like for example year C20 they uh require you to to emit event but other than that it's like up to you if you don't want to emit events you don't have to so more and more contract protocols today try to minimize the cost of gas by minimizing the number and scope of events that they uh generate another thing is you you always have to remember that events show what developer wanted to show you it's implemented in the contract code and it's emitted only when developer wanted to emit this event and it contains only information that they want to be included in this event so basically events they show you what developer wanted to show you it's not necessarily a bad bad wheel that somebody is trying to hide something but somebody may like forget about something or people may find something new interesting that was not implemented in the original contract and in such case it's very hard to do anything about it because if the event is not implemented in the contract code this event is simply not not available you cannot like add it retracted and of course events they have very limited number of parameters if they are very very complex they cost you more so people tend to emit simple events and with very limited uh number of information so again for simple things they are good enough for more complex things sooner or later you will realize that events are not enough and I would say that events are completely redundant and completely unnecessary you can do whatever you want you can analyze exactly the same things without events you don't need events to analyze everything that you today analyze using events why because actually everything is in the storage storage is the ultimate source of everything the fact that you own a token is not because somebody like aggregated historical events transfer events and realize that you should have some tokens you have tokens because it's written in the storage so it's the storage is always the ultimate source of Truth and actually this is the only thing that is important so if you have access to storage to every single change in storage you can get Absolut Ely the same information that you get from events because you don't need transfer event if you see changes in the storage if you see that in one address it was decreased in another address it was increased in the same operation you can easily get the fact that the tokens were transfer so you don't need events if you have full access to the storage but it's not that is working with storage is extremely hard when uh again when evm was designed nobody was thinking about using storage as a source of analytics this is why developers they they created events as something that's usable and easy to use because storage it's big it's hard to access actually and it's it's very hard to make an sense out of it first of all as some of you know it's not easy to get the storage information from the node because most of the storage information is a hashed in is in a hashed form if you have hashmap like balance off very popular thing this uh uses some storage locations and those storage locations that are calculated as a hases of on address and the memory slot so actually if you see hash you cannot reverse this hash and you cannot just by looking at the note see who is the owner of the token the owner for this balance that's written in this specific memory address this is why everybody knows we you cannot reverse or you cannot iterate over hashmaps right you cannot just just uh easily look at the note storage and get list of owners of tokens right you cannot get list of owners of tokens just by looking at the balance of uh hashmap in the note storage because these are hashes they cannot be reversed this is how it's designed this is one thing that makes uh working with starage very very hard uh it's very easy thing if you for example work with eer scan eer scan shows you the storage changes for every transaction you can go there and you will see like huge hashes mostly that cannot be used for anything another thing another problem with storage is this ambiguity which means that the same location in storage can have different meanings because for example uh you may have uh diamond pattern cont contract you know Diamond patter contract is a contract that has uh one proxy and multiple implementations at the same time and two different implementations may use the same storage slot and may have like completely different usage for this and different meaning for the same storage location this is not like uh very very often it's not happening quite often but it's it's possible another thing is it's huge like this blockchain story which is is really huge very uh big source of information and to analyze anything using storage it requires lots of analytical skills it's very technical lowlevel stuff so this is the reason why people mostly don't use storage for analytics this is why everybody goes to places where you find events and trying to build analytics using events because working with storage is is very very hard and uh what we do we actually try to solve the problem first thing is to make the storage information accessible to do it as I said we cannot just query the Noe because if we ask the note for history of transactions like using RPC standard way of getting history of transactions we will not get this data because it's hashed and some information is like forever lost for example if imagine that you have one transaction that changes some storage three times note doesn't store it note has only value before transaction and after transaction so note does not store does not uh provide you the information of anything that happens inside transaction I think it's even worse because it's the beginning of the block and at the end of the block so actually if you have multiple transactions changing the same storage slot if you ask the note you will never get this information this is how no this design so we created our custom Tracer which is extending the functionality of standard evm nose and with this uh custom Tracer you we actually replay every transaction and we get all the all those storage changes every single storage change including the Trent storage and we get it in preash form so we don't see that hash was changed but we see that balance of myself was changed so we see the keys that were used in hashmaps and and many other things so first thing is to get this information in preash form and second thing is get every changes so basically every uh operation that changes the the the storage uh in uh uh in smart contract storage another thing is to make it more useful for people so we also decode this information because what we get is we get information that it's slot number two for key 0x something something it was changed from A to B but what is slot to if you have contract code you can quite easily get the variable that is actually represented in slot number two so you can uh quite easily translate slot number two to balance off or to reserve zero or anything like this so you like give the meaningful names to every variable that is used in the storage if you have contract code for this we created our custom parser for contracts today uh solidity compiler it gives you the actually the storage layout but we because we need a slightly different and deeper information we actually parse ourselves every contract that's available so for any contract that we can get source code for we are decoding so like giving the meaningful names and proper types also if we know what is the variable we cannot only give it a name but we also can represent the content of this memory location in a proper proper time then we put everything in a multichain database where we integrate this with BLX transaction events calls so basically you can use all of this because every this storage change it happens in the context of specific call of specific transaction specific block so this context information might be very very interesting and then on on top of this we provide some analytical fronts but we realized pretty quickly that it solves only one problems it solves problems of availability so you can have access to those storage div in preash form decoded but still people cannot use it still it's very complex still it requires lots of analytical experience and compared to events which are very well very very good for analytics because every event is actually uh you may call it also observation because it happens at specific moment in time and has some measures some Dimensions potentially so it's very easy to aggregate them it's very easy to build queries using event we' storage is not the case because changes in storage have completely different nature and to realize that even if we provide every single change in storage and we provide this meaning like names and types it's still still very hard for people to analyze it and we started thinking about how to solve these problems of complexity and what we are working on right now is what we call EOS EOS are similar artifacts like events so they are emitted in specific moment so you can emit Eco when something happens which means that you can emit Eco when block starts when transaction starts when specific function is called when specific slot in storage is changed so you define the moment in time in execution of transaction when Eco is emitted and you define the parameters like for events which are calculated using the whole blockchain storage at this specific point in time so Eco so when you uh Define this trigger point for example emit this whenever new owner is added to the multi which means when the list of honors is changed so whenever list of honors is changed by any way I don't care how it was changed I'm just looking at storage and whenever the list of owners was changed andit this Eco which has every information that uses any piece of storage at this specific point in time so at the end of the day you get something that looks like events because they are like hooked in specific moments in execution of transactions and provide you with the list of parameters that you that you define but how it differs from Events first of all they avoid traditional limitations U because they are not onchain this is something that you calculate offchain using the data that you get from the node so if you have like database with full history of every single storage change you can calculate whatever you want off chain like using this data building some queries some calculations some formulas that use storage so basically it does not cost you nothing in terms of gas of course it cost you some computing power and and uh uh Computing this offchain of course is also uh connected with some costs but it's like completely different thing they provide you complete transparency because you get what you want it's not that somebody emitted event at one point in time it's you wanted to see what happens in this specific moment because you have access to every the whole blockchain state you can like include informations from other contracts that were not involved D in this execution you think about for example you see the transfer and you do the typical transfer event you can add the prices for example you you get the state of some pools Unis swap pools you calculate price at this exact moment and you create something that's like transfer plus prices or you generate something that's exactly like transfer but also includes the balance of the the receiver after the transfer so you see that I transferred to kishek € 10 put and it and the balance of kishek wallet after the transfer is 1,000 if right something that people need but it's very hard to get from traditional events so you can include basically anything and if you have multi-chain database you actually can in theory include also information from other blockchains because in the moment when this transfer was executed you can connect using time what was the price on other chain right you are not limited to the scope of the storage of this specific chain because you can join information from multiple chains using time it's not perfect of course but the only way to like integrate information from multiple chains but you can do it so that's what you get and some examples ah so maybe before so they are uh generated like directly from the blockchain state if you have access to full State you can transform this state because Eco is nothing else like transformation of the St transformation of storage it's a formula that makes it more useful for people we don't generate any new data we just transform existing data you can easily calculate it retroactively for the whole history some of you probably are familiar with the concept of Shadow events something that's also very interesting uh concept and is also like in the same tries to struggles to solve similar problems uh problems that some events are not available or you want to have different events with Shadow events approach you actually modify contract code you add some events that you want to be included and you replay the contract like you reexecute the whole history using evm but with modified Contra which is very costly because you know it requires like execution of the whole history and is not perfect because if you change the contract code because basically Shadow events changing changes the contract code it may interfere with the execution so just by adding new event new Shadow event you may change how the contract is executed because things change like it's another different action real action in the contract execution so we avoid all those limitations because you do it like completely retroactively just by querying and building some formulas on existing data so you define when and what and that's the definition of events at the end of the day you get just the data set that you can analyze in exactly the same way as you analyze events today so very simple examples this change of signers you have safe wallet like probably everybody uses some multi here and normally how you see who is the signer of the wallet you you look at the UI right you go to the save web page and you look at UI and you see L of signups if you want to be like do it yourself you don't trust this UI you look for new signer events okay there is something like new signer event so you look for new signer events and you or change signer or remove signer and you iterate over this history of events and you see that what's the list of events today it doesn't work if you know or you don't know you can add signer to multisig save multisig in completely quiet way in a way that it's not shown in the UI you not you you don't see it in the UI you see no event nothing but there is new signer at if you are interested look for tweet from boseski from l2b he described this once in very details using our data by the way and you can so so you can actually add this Shadow honor to multic if you use events you will never see until the signer signs something example of Eco is look at the list of signers which is just portion of storage and whenever it's changed emit new signer Eco so I don't trust developers I'm don't trust that they have not made any mistakes I don't trust any like potentially I don't care about any back doors anything like this I just look at the list of signers list of signers is a portion of storage of this multic so whenever it's changed I'm emitting Eco I don't care how it happened it happened so it cannot be hidden from me whatever you do if you change list of signers I will and me e another simple thing is like extending the transfer like the probably the most used event is transfer event but normal transfer events they lack many interesting informations and for example I'm building something and I only want to get the uh information when there was some change of balance for holders of a token and only big holders ERS of the token so whenever there is a change in balance of holder of this token regardless how it happened I don't care if it was mint burn transfer rebasing whatever I don't care I just look at the balance off whenever it's changed I generate something that has time what function was used who was sender who was recipient what was amount what was the sender transfer cost amount what was recipent balance post transfer what was the price of this token in Unis swap at the moment of transfer who was the center of transaction how much gas was used there many things that I can put together and created this analytical artifact then I can analyze it like traditionally as with any other tools that's that's what we are working on now to make it available and easy to use we have already all the storage information and all tooling to provide the storage information it's just a way to make the storage information more useful and easier to use for people that are used to traditional events so as I said this provide this transparent completeness and basically gives you ability to analyze what you want to analyze not what somebody else want wanted to show you and that's it and if you have any questions I'm almost on time how much additional storage is needed to operate uh and to be able to use EOS in comparison to like evm machine without those features it's um doesn't impact the evm machine because you know evm machine is the same evm machine is the same the only thing that we change we add this Tracer which means that if you call the debug Trace function using RPC it doesn't give you just the list of calls of traces but it also includes all the changes that were that Happ but then you need to store it somewhere so it's not impacting like the EDM and no it's impacting your storage because you need to create like your Warehouse your database where you keep it or we keep it right and provide it to you so uh compared to traditional uh I think it's like more than twice storage itself is more than twice uh all other information that is like trans transactions blocks traces events like some together uh actual number of TR hold I don't have time but if you're interested we can like look into the data and I can like give you exact numbers it's not insanely big because the data itself is is simple it's like when happens so what was the uh the the trace that actually uh made this change also we have like ordering in Trace within a trace so we know like in which order they happened which is very important and then you have the memory address and value before and value after and then decoded information which means what what was the name of variable what was the key used all this so this data is is like many rows but not very wide rows so at the end of the day it's for ethereum I I'm I'm guessing now but I think about 50 terabytes for just the the storage uh information so it's a lot but not insanely big thing and of course to query the whole history you need to have like the proper database engine to be able to do it we use snowflake for this and uh again if I have more time I can show you but you can query and bu those queries in in really reasonable time it's like up to few seconds to get information about full history of uh changes for some contracts so it's it's pretty doable thank you so much Toc thank you

Automatic transcript — names and jargon may be misspelled.