Aligning incentives to enable decentralized storage and compute - Vukasin Vukoje | Alt Labs
ETH Belgrade Community·Sat, Oct 7, 2023, 12:00 AM
Speaker
Transcript
[Applause] foreign I'm the founder of web tree mine and today I'm going to talk about how we think about decentralizing the cloud uh which in the Webster context means how do we think about aligning particular resource providers which can be storage or compute and ultimately executing a particular workload so the reason why we're working on this is the fact that we ultimately think the internet is broken today and the reason for that is that in a way we are not controlling our digital lives most of the data that we store most of the data that we generate is automatically being gifted to a particular set of organizations and we really don't have much control over that so we have no clue where out there resides so basically we don't know which server where which redundancy and so on uh then we have no idea on how decisions are made for us so if you go to the uh your Instagram feed you're just going to see a bunch of people but you have no clue how Instagram actually decided to show you that particular feed and what that means is that basically your in in a way like your digital identity is controlled by how algorithms are thinking about you and guess what most of those algorithms are not really written by people that you should trust rather are built by organizations that are as a reason why they exist are for-profit organizations and we're all fine with that and uh in a way like that's also expected from them and they are not really evil people because ultimately organizations need to somehow survive but the issue is that the fact that we don't have transparency over that and that those organizations are the ones that control most of the data that we are generating today now yeah ultimately for-profit organizations are making all the decisions related to our digital identities we have tried to tackle this like a bunch of stuff on gdpr a bunch of discussions in Congress But ultimately like all of these are kind of like trying to tackle something that is a fundamental core infrastructure problem as a more of a social problem which is not the best way of approaching it so ultimately the infrastructure is broken and the reason for that is that everything is centered around legal entities not users so like if you store a particular file that file is automatically gifted to a particular legal entity you as a user are not retaining any rights on that file in fact like your just check boxing like a bunch of stuff before you uh sign up for those services now the reason why that is the case today is not again because uh those organizations are evil or whatsoever uh the reason for that is that in the 70s the way that we design like all the paradigms that we're uh working with today were mainly meant as a way for us to have something that was relatively performant so the way it worked is that you had to figure out a way to retrieve a particular file and it made sense that you have a big machine on one side that was the server and you had like users trying to retrieve files from that big machine and that's just because it was more practical but ultimately that client server model stick and today we're still using that even though there is no reason for us to be using that so because of that of course location addressing is by default so like basically whenever you're trying to retrieve a particular piece of data what you're doing is you're looking for the location of that piece of data you're really not asking for the data you're just asking for the location and of course we have a bunch of abstractions that are helping that uh on the devtool side to actually try to like manage the data well and like we have a bunch of work that was done there but ultimately you're asking for the location and the issue with the location is that the location at the end of the day is going to be a server and the server needs to be owned by someone and that usually is a legal entity that needs to be incorporated in a particular jurisdiction and basically needs to be controlled by a particular entity board shareholders and so on and again is for profit so after that because of the way that like we designed the internet all apps were kind of built to be location addressed so like everything that is built today is by default location address and the worst thing of all is that all Dev tools that were actually built on that server client model are location address meaning that's like if you want to build anything today you need to like do it in a location uh addressing uh like paradigm now what does this particularly mean for each of us here it really means that by default all our data is not really ours so uh ultimately it cannot be ours the reason for that is that the infrastructure is broken so what we need to do is we need to rebuild the infrastructure which is a big thing since we were using something from the 70s and a bunch of work was done like with a particular mindset around like location addressing now we need to rebuild everything from scratch so and all of that is just because of the wrong Paradigm so if you go a bit deeper into like how that part of them is wrong uh ultimately again because we are referencing the by location one single entity is dominating a particular use case so is it really natural that we have one company attacking one particular use case why don't don't we have like more Facebooks like it's one company one ownership structure and a particular shareholder groups that just cares about the profit and we are fine with that no one here is using alternative to Facebook we are all using Facebook and we are also calling those use cases by the name of the company which is wrong but we are just used to that if you just go to Google you're just gonna say I'm gonna Google something why search dominated by one particular entity with a particular set of shareholders in a particular jurisdiction that should not be natural and all of that is just because this centralized infrastructure is owned by one particular set of uh like entities and those are all the cloud providers which is the reason why we are trying to disrupt that now uh ultimately what I mean all the time by like location addressing is that when you go to the URL in your browser you're always gonna have a particular location and that is something that today most of the time is kind of hidden for you like Safari is just showing you like the main domain like they don't even show you like the location of the file because like it doesn't matter anymore uh so what we actually want is we want to go back to the roots and really understand what we need as Humanity so what we care about is the content is not the location so if you want to reference something we want to reference that by the content as well what we care about is we care about having multiple organizations working on a particular use case so like we should have more Facebooks like why do we have only one and the reason for that again is the centralized nature of the infrastructure because if you have a decentralized infrastructure you could imagine having multiple apps that are doing like your social media use case and you're consuming the same data but actually that's not the case today because like someone has a monopoly on your data and it's not letting anyone else access your data which is the reason why only one Facebook exists but if we had like a decentralized infrastructure as a common good like more as a water or electricity we would be able to actually have multiple apps consuming the same data that is owned and controlled by us and then we could think about potential use cases where we are keeping that the private we are encrypting it we are allowing particular entities to see that they are if they incentivize us in a particular case maybe we want to sell the deal or whatever but then we have rolling keys and so on but that's not the case today so ultimately we should ask a network for the content and if we ask the network for the content what we're gonna get is we are going to get back the file from a particular node that can be like in this room it can be like in this country it can be in Europe or wherever but it's not controlled by one particular entity it's a common good that we can all access and depending on whether the file is encrypted or not you can see what's inside that file or not so all of this is kind of cool but it kinda seems like science fiction but actually we don't think we are that far away so if you look uh today like we have a bunch of progress that was done by like numerous organizations like super reputable super smart folks and focus on a different set of problems now uh in the case of Falcon we have a network that is focused on storage that network is great for archival today I cannot say that it's great for retrieval as well if you would wanna do retrieval you would probably use something that is not incentivized today like ipfs uh then we have life here and that is folks focused on uh encoding oh yeah like we forgot to change the library to encoding but yeah it's focused on coding renderer is focused on rendering and chain link is basically allowing you to just link a bunch of different networks and get informations that you need in order to build smart contracts or a particular logic that is living off chain so if you go a bit deeper into how these networks work so like you have Falcon which has like a group of storage providers so those are organizations that are providing storage to an open network of course you can identify like those files on that Network through content addressing so like you have cids that you ask for and you would guess like uh more information about where you're there resides and I'm going to go into that a bit later but the important part is that every 24 hours there is a proof that is a formal proof that is generated by that particular search provider uh proving that that particular piece of the is stored on the network this is on every 24 hours for every 32 gigabyte sector uh with a snark uh and that is on chin forever so you'll always know that you're there is going to be there forever in the case of like them committing of course and there is something economics in the background as well but ultimately we all think about Falcon as something that might come like uh and like of course like we all waited for the launch of Falcon for so long that ultimately today like most of the folks don't even think that focus on protocol labs are doing much but actually they are so like if you look at these set of organizations those are great organization you have certain you have Berkeley you have Solana the internet archive is basically our Hive and gold idea uh that they are generating but are having the internet every 24 hours on Falcon which is like a stupid amount of their like petabytes now if we go a bit deeper into the details on how like a storage network works and by the way I'm talking about the storage piece because the storage is the core here like because on the storage we're going to store the data and this is how like everything is going to be architected around that so ultimately the compute is going to come on top of the search so the reason why I'm uh diving a bit deeper into the storage piece is that everything else is kind of going to be similar so ultimately you're always going to store files in sectors so like the sector term is used by the drive so if you have a hard drive you have sectors on the drive so like it was a nice way of like uh fragmenting a particular search Provider by having like a bunch of sectors stored on that storage provider so sectors can be empty so if sectors are empty that means that the storage provider is storing just zeros and is generating a snark on top of those zeros and he's in a way committing capacity so that storage provider is basically just storing zeros but you know that there is like a capacity that you can use as a storage client on that particular storage provider which can be like in a particular jurisdiction or so on now the way that you've write to that sector is basically uh you do a snap the snap is also like a snark and you're doing some cryptography and basically you're replacing the zeros with a bunch of uh there so like you're just taking a big blob of 32 gigabytes and you're just snapping that into an empty sector uh ultimately one sector can contain multiple files so you have a bunch of files and you have the sector so if you want to retrieve a particular file uh you will ask that storage provider whether that search provider has like a particular CID which is the hash uh you will provide the hash the search provider is going to tell you yes this CID is in this particular sector if you want to read it you need to get the entire sector now this is as well why I'm saying Falcon is great for I hover right now not for retrieval yet but ultimately with the tech maturing is going to get better and better because you are transferring 32 gigabytes of data just to retrieve maybe a couple of megabytes of a file ah so ultimately like all of this is off-chain State because like you are storing something off-chain there is a proof that that particular piece of the is off chain but ultimately it is option it's not something that you can read right away nor you have like knowledge on whether that particular state is on on chain in that particular time of course that is very different from like chain state where you're basically just storing the same state over a group of providers uh yeah in the case of ethereum manners however we want to call them but basically you are not storing like the same piece of their on every provider you're storing one sector in one single search profile now what we do actually is we just connect all these networks because like all of those networks right now are completely independent like completely running on a different set of storage providers compute providers which is a big issue right because ultimately at the end of the day if you want to consume a particular piece of the air like you need that to be at least in the same local network you cannot be retrieving there like from all around the globe which is not something that is kind of natural for us in the context of crypto like we were usually focused on like just chain State and we don't care where people are but ultimately when you're moving a lot of there back and forth you really care where the data is because you are limited by the uh bandwidth that you have across the network so the way it works is that we basically use Liquid sticking to incentivize a particular group of computer providers or search providers to store a particular piece of data or to do a particular thing so like if we could say uh we are going to allow you to use this liquidity that we have only if you run also compute uh like infrastructure in the same Data Center and ultimately you'll have compute providers and storage providers that are kind of having some overlap but you'll also have situations where you have a very specialized compute provider if the computer job that that compute provider needs to do is a very like high compute intensity workload so for example if you're doing rendering or encoding that doesn't make sense to potentially be done in the data center where they are resized but you want it to be somewhere close you want a couple of milliseconds of latency uh so in the case of compute providers you can imagine of those folks uh basically as folks that are doing a lot of encoding rendering so have a big GPU farmed and so on and the way that we would have the communication is by having like a bunch of like storage providers that are close to a particular set of compute providers and you want a lot of bandwidth to be available between those two uh the computer providers are more like ethereal manners where you have a bunch of people that have like a bunch of gpus uh why this is important because we want to commodize the compute resource so like uh ethereum manners in general GPU mining really commodized the GPU Resource as a thing so like if you look at the profitability that ethereum miners had it was always close to like zero margins which is what we want we want the compute to have the cost as it would uh be like energy plus the Opex that it takes to actually maintain a particular set of Hardware plus some amortization of the capex that was invested which is by the way not the case with Cloud providers Cloud providers are taking 20 30X margins which is a big issue by itself and I'm not going to get deeper into that well storage providers are basically folks that are just storing a bunch of data in order to store and or have a bunch of data there need to be super efficient on how uh they sort the data and usually Rackspace is super expensive so you can see a bunch of storage providers having like very dense apps so they have like a bunch of hard drives in a Data Center and like they have super lean infrastructure and they're just basically petabytes over there like in one like rack so yeah ultimately at the end of the day what you're going to have like we are going to have like those compute providers and storage providers somehow like talk to each other right now the main objective that we have is just basically having the same workloads run on a uh overlap of compute and search providers because otherwise the compute jobs are not going to be able to consume uh the the search and uh and that's the first step again like if you look at how long it took for the location addressed internet to evolve uh this is going to take at least one third of that and if that's one third that's probably a decade so we want to be just aligning everyone and allowing folks to actually run compute workloads where in the future you'll be able to execute uh maybe the encoding of a video and then have some decentralized filter that basically allows you just to remove like the background or do something else on a stream that you're basically owning without having anything to do with centralized providers and for a fraction of the cost like probably one twentieth so yeah I think this is all I had yeah like I'm going a bit deeper here into the fact that we are L2 but this doesn't really matter for for these audience so thank you so much uh it was great being here and especially taking the last lot so yeah just grab me for a beer or something if you are interested in this [Applause] if we have any questions for Booker feel free to raise your hand yes over there in the back do we oh um so you mentioned basically that you're using LSD derivatives to secure the network does it make sense to maybe use taking a falcoin you know because like LSD obviously kind of gives you a derivative you can trade or whatever so don't you kind of lose security because of that so let me try to reframe and see whether I understood the question so you were saying that we have LEDs so instead of having Galaxy is like would it make sense to use the native staking was that the question yeah there is no natives taking like that's also something that probably I should have covered we had to invent staking ourselves so the way that the Falcon economy works is uh that in order to provide storage you need to also have a bunch of liquidity that you put as collateral so like if that the particular piece of the air goes offline the storage provider is getting slashed so they are losing the tokens so ultimately uh the way that Falcon was architected is that with more search joining the network more liquidity is going to get lost locked and uh that's kind of how it works so the way we approached it is we were like okay so this is a nice uh way for us to actually have leverage and make sure that the storage providers that are using our liquidity are actually doing what we think is right for the ecosystem as well and on the other side we have stickers that trust us that we're gonna make great decisions for the ecosystem and this is why they are sticking into our pool at the same time we need to make sure that those search providers are super efficient and if they are efficient we are able to like pass back most of the rewards that are generated by uh the search back to the stickers so ultimately there is no staking all the block rewards go to the search providers but the liquidity that you need to lock is super high and ultimately you need to like provide that utility as well which is storage but the liquidity is much more scarce so like if you have one rack full of drives which is probably around yeah it's around nine petabytes if you store Nine Petals of there like you would need to have probably 50 million dollars of liquidity that you need to lock in that rack does that answer your question do we have any more questions over there okay so I was wondering like basically in in a centralized world a world when you like uh let's say you have a bigger image you put it on let's say a S3 bucket then you have a cloud front that edge is that and then you basically have like a retrieval time that's minimal so I was wondering like when we are looking at decentralized storage uh what's the user experience at the side of fetching things basically yeah that that's a great question so ultimately if I understood something in the past I don't know three years at this point working in the Central Storage is that uh even all the things that we take as granted that we have from the cloud we're actually slowly built and you have so many abstractions that you're not even aware of so like even AWS is not really building everything themselves so they have first they have like some kind of rate on the hardware level usually that's done by Seagate or something like that then you have a bunch of redundancy so like they're going to have if uh like if you take like the cheapest one they're gonna have like uh redundancy two so they are going to keep one particular sector so a particular piece of data into locations then you have like Geo redundancy and then slowly they are just increasing the the amount of redundancy and they just to Source something if you would want to retrieve something in the context of web tree uh you're just going to do that through ipfs but even there you need to have a set of nodes that need to be super fast and you'll probably use something like web3.org nft.sorge which is basically a centralized provider of that particular file that has a particular incentive to provide your CID back super fast and slowly we are going towards the mall where even the providers of that retrieval layer are going to be incentivized to give you back something super fast and you have initiatives like Saturn Network that are basically trying to do that but ultimately the way it works is that on one side you have Falcon which is just for archival you're not gonna even touch that oops uh you're not even going to touch that and then you have ipfs and on the ipfs side you're going to have a bunch of tooling that will give you back the file super fast and usually that is going to be on a CDN or something like that um can you can you get the microphone please we have time for one more quick question before we move for the conference closing does anybody else want to score uh so that sounds a bit centralized at least in the moment I mean we're having this decentralized part where it's archived but then we actually have a centralized part that's caching that side and offering you two two games depends depends on how you get back the information on which node your data is so you're gonna say I need dcid and you'll have like a set of indexers that know like on which storage are heavily providers and retrieval providers your particular file is and you're gonna get like the list of providers and you can decide where to retrieve it and all of that logic can just sit on your client and those indexers can be like super decentralized so it's just about like decoupling uh the responsibility there and also relying on the tag that we have built ctns are great like there is so much work that was invested in the end we should not try not to use cdns but we should try to control like which cdns we are using and basically treating them like as infrastructure providers and not to someone that can own her there okay okay thanks and with that being said it's my honor and out of its sadness to say that this was the last talk of it Belgrade
Automatic transcript — names and jargon may be misspelled.