TrueBlocks Workshop | Thomas Jay Rush | ETHDam 2023
CryptoCanal·Sat, Oct 7, 2023, 12:00 AM
Speaker
Thomas Jay Rush is a president of TrueBlocks. https://twitter.com/tjayrush?s=11&t=3zlJQECNQOGVgRU84E-vDA TrueBlocks https://trueblocks.io/ ETHDam is a Hackathon & Conference that gathered over 500 DeFi and Privacy builders on the 20th and 21st of May 2023 in Amsterdam. Privacy is normal. Following the arrest of Alex Pertsev, a Tornado Cash developer in the Netherlands, ETHDam 2023 is determined to counter the chilling effects of the lawsuit and bridge worlds to discuss the future of privacy and encourage to build on the shoulders of cypherpunk giants. ETHDam is powered by CryptoCanal, - a blockchain education and events platform growing in Amsterdam, spreading its roots to Rotterdam and Zurich. ETHDam 2024 is on the map already! Keep up with us to see updates: CryptoCanal https://www.cryptocanal.org/ CryptoCanal Twitter https://twitter.com/CryptoCanal Join CryptoCanal Community https://t.me/CryptoCanalCommunity We would like to thank our partners and sponsors that made this event possible. 🌷 Our BFF 1inch https://1inch.io/ Our Frens: Sismo https://www.sismo.io/ Aleph Zero https://alephzero.org/ Scroll https://scroll.io/ RAILGUN https://railgun.org/#/ And our Sisters: oasis.app https://oasis.app/#earn Maven11 https://www.maven11.com/ bitvavo https://bitvavo.com/en Lido https://lido.fi/ Spankchain https://spankchain.com/ API3 https://api3.org/ Gelato https://www.gelato.network/ VanEck https://www.vaneck.com/nl/en/crypto-etn Marlin Protocol https://www.marlin.org/ Silent Protocol https://www.silentprotocol.org/ Cyber Capital https://cyber.capital/ … and Proto https://twitter.com/protolambda 🍍
Transcript
foreign welcome maybe we can move a little closer it's fine my name is my name is Jay rush I'm going to talk about true blocks which is a project that I've been building for about three years the fundamental idea is that I think that if you truly want to be Sovereign and private you have to have a local source of data you can't be reliant on some third-party provider for data so for me the easiest way to do that is to run the node software locally on my machine so this laptop here is running in this screen it's running prism and this screen is running Aragon so I'm going to start both of those pieces of software up so that's prism and this is Aragon I'm going to do something weird here I'm going to turn off my Wi-Fi because I want to demonstrate that this is a truly local piece of software and it's truly decentralized Aragon starts complaining like crazy because it's saying I can't find the internet but as soon as I turn it back on Aragon's fine so I want to build a piece of software that has a purely local source of data and has no Reliance whatsoever on third parties so that's what true blocks does so I'm going to switch um I'm going to switch to another screen and talk about true blocks I'm going to give you a little view of the server in the background this thing is still complaining like crazy because it doesn't have an internet connection but true but the node still serves data to the application so we have an application called shifra which is part of true blobs and it has a whole bunch of tools one of which is blocks so I'm going to say she from blocks 12 and it's going to deliver to me block 12. I can say sheep for blocks 12 to 10 million in steps of a thousand it delivers all the blocks and and every one thousandth block from the chain totally disconnected from the internet which is exactly the point of what I'm trying to build so you could use this for example for data science um one of the really nice features of this blocks command is a is a sub command called unique and what that does it lists all unique addresses in a block so I can say block um one million and one that's all the unique addresses in a block and I wanted to build that because I wanted to build an index of addresses and where they appear in which block and in which transaction so I can do a more complicated block 17 million so there's hundreds of unique addresses in Block 17 million one hundred thousand so but we want to know where every one of those addresses appear because we're building an index because in order to make the node software that's running locally effective as an application server we needed an index we needed to be able to index addresses so um let me close this as I said shifra has a bunch of different tools one of them is called demon which starts a server and serves all of the other tools so I'm going to start that so here I started a server that's pulling data only from the node and it's serving all the commands for shifra and then we build applications on top of that so this is an application it's running locally I'm still disconnected from the internet takes a while to start up but this is our application that's pulling a history of this account from directly from the from the directly from the node software and this application is really really interesting to me because it has all these qualities that you wouldn't expect it's really fast I just look through the hist the the 5 000 transaction history of an account which you can do if you go to a web server but you're going to get rate limited on that web server you're going to get one page at a time and there's a reason for that that's because the web server is being shared by thousands of other users this is one user for one application I can hit my own web server thousand times faster than I can hit a remote web server so I can get a lot of data into an application way more quickly and we have now an index of where to look so that index is really important for the application because I can query and get a very precise list of the transactions that I'm interested in as opposed to the way the node software Works where I have to query a range of blocks and look for data so this application is rough I admit I'm not a front-end developer but it has these qualities that are that that are really strange there's no pagination I can scroll through this like I'm I'm scrolling through a through a spreadsheet and what we do because we have every single transaction is we can do what I call a Reconciliation so here this is saying I spent die I had 2 000 die I spent 90 and now I should have 2 000 you know 90 less than two thousand and it reconciles and then here I'm going from I know I had this much diet the last transaction so I have that much diet the next transaction so we can reconcile our history perfectly throughout the entire history of an account and if you have that you have automated accounting and to me automated accounting off chain if you know anything about trying to do that you know that it's nearly impossible the reason it's nearly impossible is because they don't have good indexing on these remote apis so I'm going to switch now completely that was just sort of a demonstration of what we built I'm going to try to help you understand how we built this so uh and this has this has another really interesting aspect to it we use ipfs to share this index and because we can do that we can make the cost of running our system we kind of push the cost off onto the end users and running the system becomes zero cost for us all we do is provide the software and I think this challenge is um any application that's going to charge me 250 a month to access tax information or transactional history information so we can wildly lower the cost of these applications and I think that's a direct result of this web 3 technology is one of the results if we use it correctly is that the cost of doing these operations becomes near zero so I want to try to get to that point about myself um I used to work at IBM research a long time I was heavily into the early internet got totally disillusioned became a poet and a writing teacher and then um came back into the tech field in 2015 when I first heard of ethereum and I've been here ever since um as I said I want to build an application that has huh I'm gonna are we good I want to build an application that has only one source of data which is the node and no third-party web apis because I think this actually delivers on all of the original visions such as self-sovereignty perfect accounting which builds truly transparent organizations if you can do accounting from the outside and that leads to coordination I talked about shifra and this is that command that gives me all the unique blocks and this is the data that we look through one of the things that true blocks does that a lot of other things don't do is we look at all the obvious places where there are addresses in the data all these pink things but then we also go into the input data and the trace output data and the log data and we scan through that those bytes and we pull off addresses that other people don't find and that's why we can do reconciliation and that's why it works um I won't go into exactly what we do but a very simple example is right here that log topic is 32 bytes long and it has six or 16 leading zeros and what we do is we say that looks a lot like an address so we're going to call it an address for our index a lot of people wouldn't include that in their index and that's why they're missing data I think we compared ourselves to etherscan we did 5000 addresses and on average we find 13 more transactions almost all of them in that input data that we looked at and then we convert those into why were they missed and a lot of these transactions that are being missed by regular indexers are Financial related functions and that's why reconciliation doesn't work on most platforms because they're missing Financial transactions so I'm just going to skip ahead here so I said we we look at every block and we build a index we build a list of we look at every block and we build a list of addresses that appear in those blocks and then a regular indexer takes that list and continually sorts that list by address because that's what an index is but we wanted to share this on ipfs and if you're continually adding new addresses into this list this ipfs address changes at every block so that didn't work for us so what we did is we basically stopped adding new data to this list and we created a um we created an immutable unchangeable Chunk we call this a chunk of the index so I say we created a Time ordered log of an index of a Time ordered log and if you think about it this is exactly what blockchains do they create time ordered logs of transactions and for the same exact reason we did this because we wanted a hash that would stay permanent forever so so we collect two million of these records and then we create a chunk and now we can put this in ipfs as a frozen forever immutable block or Chunk we call it that whole process is called scraping and we a normal API a remote web server takes this index and delivers it directly to the user but you can see the relationship here who has the data the indexer has the data so this end user is completely at the at the control of this data provider they can cut off the access they can change the data they can lie about the data they can withhold the data they can charge for the data so that that to us was not web three and we want to deliver the data from an immutable store which allows us to just publish it once and never kind of have to publish it again so we move the end user over here we write our data to ipfs with a hash an ipfs hash and then we keep a list of all of these hashes in a thing we call manifest and we publish that manifest ipfs as well and now the end user can get this manifest and in the Manifest is every part of the entire index and we publish it to ipfs so we can't take it back if they have the hash and they acquire it they have it forever and we can't take it back we go further to publish this hash into a smart contract that we call the Unchained index and then the end user can use these command line tools on a local computer to basically download the entire index if you want and then he can use other tools to serve data to his application the trouble with this is 100 gigabytes big this index so there's a really beautiful data structure called a bloom filter which lets you basically say is my address in this in this list and this thing goes from 40 megabytes to like 40K or something so I get a bloom filter associated with each one of these chunks as well and that gets added to the Manifest as well so when the user downloads he gets three gigabytes so the user can download an index of the entire ethereum mainnet chain for three gigabytes and he can ask questions can you give me a list of every every appearance that my address has ever appeared in by querying this three gigabyte Bloom filters getting a result that says I have two or three of these chunks and then downloading the larger chunks and now he has this very fast very detailed view into the ethereum mainnet completely unrelated to a publisher there's nobody withholding the data here I can withhold publishing the hash but this is permissionless so anyone can publish this hash to it it's purposefully been permissionless so what I want to see people do is realize that I'm publishing a hash to an index to every single appearance of every address on the Chain once a day or once a week anyone can pull that thing and have an index for themselves and build local applications and I did that on purpose because I wanted to destroy I want to destroy the ability of these people to control your access to this index and I I don't want it to be me doing this I want you to run this and published ipfs and share on the smart contract so we get hundreds of people publishing where the data is and we've destroyed the ability of these guys to charge us 250 dollars for it a month the other really beautiful thing that happens hard to explain is you get that chunk this chunk and that chunk I get three or 10 or 50 other chunks and what happens is I pin all of these chunks down on my own machine and the software automatically pins the chunks that you download on your machine and pretty soon we're all pinning our own data and we're all sharing it with each other so we get this we get this naturally sharded database and another thing that's really really interesting is I'm a small user I have a 5 000 transactions so I'm going to get 1 100th of the entire index uniswap if they were running a system like this would get every single part of the index and they would pin every single part of the index because they appear in every single part of the index and that's to me Fair unit swap should be sharing the entire index with the whole Community they're using 80 you know 10 percent of the entire resources of the system so each person who participates carries a burden that's exactly equivalent to their usage which to me is fair so I think that's a very interesting thing that just falls out of there's no design decision to make there it just falls out of this way of doing it and I think this is like a web three-way of doing it and today what we're building is this and we're literally gonna we're literally walking into a buzz saw in my opinion because all we're doing we're giving away all of our sovereignty to the people that have the index which I think is just a bad choice in my opinion so uh just to conclude I think I must be close to dania um I've taken care of Webb I've taken advantage of web three I didn't run away from web3 I didn't use web 2 solutions to try to build the solution I wholeheartedly embraced web3 the index is completely uncapturable once once it's here I can't stop anyone on the planet from looking at it and getting every one of these chunks as long as someone is pinning them I think maybe the ethereum foundation should pin them or some but something or ipfs should pin them or something it's uncapturable though because I can't intercede between the user's local machine and a smart contract I can't say that he can't have access to that um It's Perfectly private it's to me the I'm not connected to the internet right now and my application works perfectly well and the node software isn't getting any fresh data but that's my choice I can choose to join the network whereas if you're going to a website for data you can't choose to join it has this other really other interesting aspect if a website gets a hundred new users in a day they have to go out to the store and buy another computer because they have to serve a hundred more users whereas if this system gets 100 new users we're getting a hundred new people sharing the index so the system actually gets more robust more distributed and more resilient the more people who come and this is on purpose not from True blocks but from ipfs That's How ipfs works um I wanted to build it so that users didn't have to do anything all they have to do is start an application the application looks at the Smart contract gets the data downloads the thing pins everything and the user is just using this really fast really accurate really deep insight into their data they don't have to do anything special it's Equitable which means heavy users are carry a heavy burden and light users carry a light burden um we're reading directly from the blockchain which has hashed data which is undeniably contains the bytes that it came from and we're using those bytes as our only source of information so the results that we get are perfectly reproducible because we come from bytes that are known to be true we do a very simple algorithm and you can reproduce it so we don't have to build some complicated prover some complicated Watcher fisherman system that has a coin on top of it to incentivize people to make this accurate it's just accurate because we're coming straight from block data and we're indexing with a known algorithm and the cost is zero I published to a Smart contract cost me three dollars a day three dollars a day to publish my entire infrastructure cost me three dollars a day I run ipfs locally I am you know people use the software and share the data they're carrying the burden and that's by Design and on purpose so um I hope I made my point um the point being that we should turn we should we should build applications I know I know this is insane I'm not crazy but we should run our own node software and we should not allow the node software to become a piece of software that serves Big Data providers because it's going to become increasingly more difficult to run a node if we allow the node to get pulled towards Big Data providers we should be pulling the node back down so it runs on laptops so we can access data without asking permission from anyone if we do that we're going to quickly learn we need an index and true blocks provides one method to do that is not the only way but I want to build truly distributed perfectly private and local first applications and that's what true blocks is about so thanks for listening and uh I'm happy to answer his questions if I'm allowed one question any question sure so with thank you so now I I see you do you rely on ipfs as a storage layer for indexing but uh is your approach storage layer agnostics so I mean if somebody would like to use some other solution for that yes storage agnostic okay yeah I should use I shouldn't necessarily use ipfs what I should use is content addressable storage I don't care where somewhere we put immutable data and to me that's content addressable ipfs is kind of the most likely place and ethereum mainnet is kind of the most likely place so yeah I'm a little focused on those but yeah I should broaden that yeah thank you Jay thanks [Applause]
Automatic transcript — names and jargon may be misspelled.