ETHWarsaw 2023: Dawid Szlachta, TrueBlocks - The index that scales. Private Etherscan in your pocket
ETH Warsaw·Mon, Oct 7, 2024, 12:00 AM
Speaker
Presentation - A brief talk by Dawid Szlachta from TrueBlocks. Learn how TrueBlocks is streamlining the process of obtaining transparent and accessible blockchain data. Follow us for more updates: https://twitter.com/ETHWarsaw
Transcript
welcome David thanks okay hi everyone my name is David scha and I'm a lead developer at TR blocks an open source project started in 2016 by J Rash where we develop an indexer for ethereum and ethereum compatible chains and today I'd like to talk about indexing ethereum which means finding all the transactions that have ever happened um together with addresses linked to these transactions how we can decentralize our index at true blocks why we think it scales better and um web free way of building software so let's start by figuring out what an indexer is um if you log to your uh bank account if you have any of course uh you will probably see transaction history somewhere list of your transaction this pretty standard in traditional Finance World however it's um bit more complicated in etherium and the reason is it was an early designed Choice made so ethereum RPC has no endpoint that we can ask for all of the transactions all all transactions of given account and that's why we need additional software which is an indexer so indexer goes back in time it goes back to the very first block ever produced on the given chain and from there it travels to the fut to the to the present sorry towards latest block so it goes block by block and for each block it finds um transaction IDs and tries to extract account addresses so if you run any indexer you will end up with the index um but where should we store it right so one obvious choice would be cloud or maybe my own server but that actually leads us to a very interesting problem because the Ser the the server owner can change the database there and they don't have to tell anyone about it so it's not transparent or they could if they wanted uh Power the server off right so we are now in a situation where um users have to trust a person or organization they don't know and I hope that you agree with me that that's not a very good idea um and also I would say that such an index stored on someone else's server is a is very hard to validate or even Impossible on the other hand um we have server owner who has to cover Network fees right now and they start charging users they they will charge users for data which is free and publicly available on the Chain so what's the solution well at trbl we don't want to be the owners of the data we don't want to be Gatekeepers between our users and the data so in order to to do it um we can't simply put index on our own server we need de centralized storage well anyone knows about any decentralized Storage Solutions out there yeah great yeah so um with ipfs personal uploading the file doesn't own the servers and also data on IP adds is immutable and because files on ipfs are addressed by their content uh which means that whenever you you can download the file from ipfs change the content and reupload it but it will get a different identifier that means that if users ask for something they should get what they want and this gives us a kind of um automatic validation so it's all cool all great but the index can be a big file and right now um for ethereum mayet it's about 140 gabt and you know we don't want our users to have to additionally download 140 gabt after downloading our software right um it's problematic so what we do is we slice the index into smaller chunks which are easier to download and instead of having one giant file on ipfs we now have um a number of smaller files and we have to keep a list of these files so we create additional file called the Manifest which is simply a list of all chunks all index chunks deployed to ipfs but users would still have to download all the chunks to check which chunks holds the address the user is interested in so additionally we create Bloom filters and I don't want to go into much detail here you can think of a bloom filter a binary small binary file which can tell us if an address is present in the given chunk but without the need of downloading this chunk and as I said they are small so so for the user it's easy to download Bloom filters it's not a problem to store them on the hard drive and for our tools is efficient to to query them so it gets even better imagine we have two different users okay one user is a data scientist and they need lots of data to do their job another person is only interested in their own transactions so they are very different different use cases they both have very different index chunks but at some point they will be sharing chunks with each other and this is huge because it gives us sharing economy which has this lovely definition but simply put sharing economy is receiving something but also giving something back so by using the technology you're helping the community and it's not something you get um in web 2 centralized world yeah so great but there's still one big problem so we're the ones publishing the index in a trbl we don't want to force users to trust us so the simple solution to that problem is to allow anyone to publish a index and we do it by storing manifest location so the list of all index chunks there are storing this manifest in a smart contract deployed to ethereum mainnet and we allow anyone to publish their own manifest and add this information to our smart contract so now users they can choose who they trust and they can even choose to not trust anyone at all because they can build the index locally on their computer and keep it there we call our index Unchained index and all the tools we are developing um are open source so you don't have to really trust us TR blos as an organization you can download the source code you can audit it and make sure it does what it should do and then compile it from scratch furthermore if you go to to our website you will find a file called um unchain index specification there and it has instructions on how to read and create the index so even if at some point we decided to stop maintaining true blocks you can still um write a very simple program that would read indexes already published or write another very simple program that that that would allow you to create your own index using proven technology but the one the technology that we don't own so that's why we say that our Unchained index is truly permissionless another interesting thing is that with ipfs the more users you get the better because the more users the easier it is to to fetch chunks well in Web Two World the more user you get the higher Network fees you have to cover and yes sure people can still make profit of course but costs will be higher with unchain index for um index publisher the cost of infrastructure is negligible and it's always free to read the index for the users so we have an index that decentralized private permissionless and cost effective but the big question is can we build on it and of course we can because otherwise I wouldn't be here right so please meet shifra shifra is our common line tool to query the chain nonchain index is foundation for shifra shifra as a tool has many um features but one important feature is an ability to to give you um all transactions of a particular account and here's how it works so the user asks for a specific address then shifra scans Bloom filters to learn which chunks it should download it downloads only a small subset of the index only the chunks that it needs then it fetches transaction details and other useful data from the user's own no and it's actually a feature that we require our user to have their nodes um and because blockchains are immutable we can cat the data and we can make future queries way way faster so I would say that we can take the same logic that we've implemented in shifra and put it in a mobile app and you know mobiles are um specific development environment because you have limited resources but the fact is that well we are resource friendly we only download what we have to download and we catch the data now if you think about it um for example e scan it's a Blog Explorer right but from our from from my own experience most people use it as a way to get very transaction history so by taking the logic it's proven to work in shifra and please don't trust me it works you you can test by yourselves by downloading shifra by taking this logic and putting in mobile app we can have a eers scan leg app that's private and decentralized and my point here is that there's a different way to build software in blockchain space um it's web freeway and it's it can be challenging because it's very different from what we know from what from web 2 and I I know that because I used to be a front end developer before I joined TR blocks so this new way requires learning learning new technologies learning protocols learning skills how to do stuff it even can require educating users for example that they should run their own nodes but it's totally worth it because it also creates new possibilities for example you get sharing economy for free in Europe or you get um free access to chain history so while there are centralized um services in blockchain space there's very good reason why they are in our space and they solve some problems um it doesn't mean that also St Rec create has to be centralized so maybe next time you see to design your next app um maybe you can make it decentralized okay thanks um any questions okay let me come here thank you for a good presentation and uh this uh good tool so uh my question would be about uh what kind of data you can cury uh with Cher at the moment or what kind of data you index is like uh all the transactions that touch the account or can you like look in inside token transfers nft transfers and so on so what what's there that you can use today we try to to index all of the transactions um and then we fetch the details from the note so yeah token transfers um we we we even had this weird case where we tried to um compile you know like um say tax record for an account and the numbers didn't add up so we went you know to the low level uh and we started finding things like for example contract sub destruct so we really try our best uh to index everything that happens is um am I answering your question yeah yeah we do uh thank you for presentation too I I would like to ask you uh is user that is contributing to the whole project as uh the one that is basically I guess paying for uh putting the uh indexes on the network is he in any way um does he earn anything because in the blockchain uh at for example in ethereum if you are the node and if you're contributing to uh basically working on of the whole network uh then you will get uh a little bit of a reward in this case ethereum how it works in this uh this project so we don't reward our users with a token we want users to share index um because they believe uh in the centralization and because they want to help each other so for example um maybe someone's working for Dow in data science team say right and they discover that that true blogs and shifra is useful tool so then it would be great if they just created an index and share it with their colleagues so their colleagues don't have to spend time building their own indexes so it's it's kind of like we are public good as a project and we want people to see this you know the the good part in public goods hi uh the bottleneck in uh on chain analytics is often not the row data but uh the labeled uh one uh so uh my question is uh is true blocks somehow addressing this this part or uh is it just focused on uh raow indexes or is it possible for you users to publish and reach the data like uh I don't know connected wallets or uh all right um so we focus only on indexing and well unch index is one part that we do the second is shifra which is tool and um so we give you a method to create index and also we try to give you best tool to get the most um important data so like transaction list or or logs or traces or um costart contract that's not related to the index but Shi can does do do it but we don't go like upper level yeah okay I think that there was one more question uh okay so so thanks for the presentation uh really nice uh tool so my question is is it uh ethereum only or you're prepared for including another blockchains so it's only for ethereum and ethereum compatible chains like noses for example yeah and I have second question uh so let's say I have an account and I want the true blocks to to track it um how does it work like is it constantly checking the node or is it subscribing how does it work yeah so um there's another tool and I didn't want to mention it to not go to detail but um it's called scraper and it constantly monitors the chain whenever there's new block produced again it's it's a it tries to find transactions and the addresses Okay so there is uh a few more uh I'll go first here we have some time so so uh I have a question first of all how can we trust uh the data you provide in first place because uh in the beginning of the presentation you have mentioned the you don't want to trust like the storage in cloud or private servers okay that's fine uh but how how can we trust your data uh in the first place because if you are making a scraper that goes from the genis block up to uh current block so you can make a fraud uh meanwhile so that's my first question yeah um that's a very good question and um so you can generate the index by yourself and you can have a friend generate index on their own computer you can compare okay but if I'm correct we are using uh the data provided by you or or am I creating the index uh yourself yeah by myself but from the blockchain or from the data from the blockchain ah from the blockchain okay index is always created from the blockchain so that's why we require node access and uh we as TR blocks we publish the index by ourselves too because sometimes people want to test true blocks and we don't want them to have to go from block zero and wait some time you know so we do it but we don't even do it too often on purpose because we don't want to be Gatekeepers right right okay thanks that's cool and the second question if I'm uh correct I can only query uh the data provide um connected with a specified wallet or can I you hold the data okay I mean you you it would be good if you know the address to start right but um there's no no limits you can use trbl to query your data or DA's data someone else's data it's you know okay so yeah yeah but that all the input is always as uh contract or wallet address I see yeah it can be both okay cool thank you thanks so we have one more question uh or maybe two uh I'll I'll give you first the mic because you you had a chance but you'll have more uh so you've mentioned that there is a contract where we store uh which data is trusted uh so so how much data do we actually need to store in such contract we only store ipfs CID for manifest and then the Manifest tells us where the rest of the data is and is also on ipfs but uh this manifest does it contain like each entry for each blog transaction what's the granularity No No so um our index is you can imagine a database base sorry and you have two two columns the first one is transaction ID the second one would uh have all the addresses involved in this transactions that's the Manifest okay thanks so it's relatively big right excuse me it's uh the the amount of data there is it's relatively big isn't it yeah it's not that big I mean so it depends on the Chain uh but the data we store on uh on the chain is small it's really it's just CID single CID and the index can be like couple of gigabytes big uh Bloom filters are for eum Min right right now it's like um three gigabytes okay thanks and it's really interesting tool I'm going to check it out thanks awesome thanks so the final question I don't remember if you mentioned that but could you tell if there's any bigger third party project that is based on those indexes did you know uh about some things like that not at the moment I don't think so um however um for example GV and gitcoin Di uh some of the teams they use true blocks okay thank you very much what's between well if you run your own scraper there's there should be no latency at all yeah if you use our index they're going to be latency because we on purpose we don't publish index too often but I don't want to use excuse I I didn't get it uh sorry so maybe the question you can ask later on cat me later uh yeah so uh David thank you very much very interesting talk and uh we'll have Round of Applause for David of Cs
Automatic transcript — names and jargon may be misspelled.