# ETHWarsaw 2023: Polina Aladina, OnMachina - New decentralised storage coming to town

- Channel: [ETH Warsaw](https://streameth.org/eth-warsaw)
- Date: 2024-10-07
- Duration: 28:38
- Watch: https://streameth.org/watch/yt-2poYw46wAbE
- YouTube: https://www.youtube.com/watch?v=2poYw46wAbE

## Description

New decentralised storage coming to town - A presentation by Polina Aladina from OnMachina exploring the advancements and the future of decentralized storage solutions. 

Follow us for more updates: https://twitter.com/ETHWarsaw

## Transcript

hey everyone uh yeah I'm about to start I'm actually very surprised there's so much people showed up I uh yeah so I was wondering like do you come here because you actually need to use your decentralized storage for something is there anyone or maybe you're building a product like that no just just free time on Friday well thank you for showing up anyways uh right so um I want to tell a little bit about myself my name is Palina I'm in this space for a really long while now uh pretty much since 2015 on bitcoin and Ico wave and whatever so yeah you should definitely listen to my opinions here and what we're going to talk today about is why we're building a new storage uh what's the point of building it because there are other storages is there even a place for new decentralized storage so here's a brief overview on uh why so uh I think well that quote kind of looks like from '90s or something but that was actually sa by Andy Jesse last year I think in May which is 95% of the world it spend is on premises and non in the cloud and we'll still think wow that's a lot but frankly what he tried to say is you know Amazon can grow to any more time from this because you know like even a 5% uh it's I mean it's still a huge oligopoly and this is a graph by a16z also made like last year or something analysis on who owns how many percentage of the cloud space and you can clearly see like Amazon is leading this thing then all the other fancy names like Microsoft and Google Alibaba they're on the bottom you don't see really anything decentralized here because I don't think there's like act decentralized cloud provider as of now that's generally used and then there is this thin line of others and what this is is all these providers that actually are one by one they run open stack for example or a ky kubernetes or Linux and they provide their services locally in France or Germany and so on and the point here is like uh what they offer to the market is not very compatible to Amazon because I mean together they're number two provider but they're kind of competing with each other they use open source most probably uh some somehow tinkered but not their own implementation like Amazon and they have very fragmented developer ux and apis are different while in Amazon it's like you just choose Amazon everybody knows Amazon right but for any if you want to use something else you always have this uh like barrier to actually use it and they don't really have interoperability between each other like the best you can do right now in the market in in the space is try to get that the same Amazon interface on top of your implementation so it's really easy to shift from Amazon to your product name the product so tldr there is a clear economic coordination failure and yeah unable to clearly nicely leverage uh combined scale and yeah um the market power of oligopoly is such that why the oligopoly generally is a problem because oligopoly uh creates a barrier to entry for new players because it's very hard to compete with one huge scale thing and then uh as a result you have all the customers because there is no other place on the market then you can just shift up the prices and cause higher pricing and that's pretty much the problem with the Amazon because they e in their own customer margins and you can see from this graph here I don't remember where take it from but uh this is how much out of the revenue these companies spent on their cloud and that's a lot and there is like actually there is a every devops has this saying that if you want to spend infinite money in a very short time you just use Amazon this is just what I found in like two minutes if you just Google search for that because yeah you just I know shift spend another instance forget to turn it off and your build just Skyrocket and so on and so on or if your startup actually becomes used by a lot of people you just have an unlimited bill on Amazon and out of all the cloud that exist uh S3 is particularly profitable and as you can see it's about estimated 50 to 70% of the whole Amazon Revenue which is I mean more than a he just for the storage and overall the storage needs are growing based on the different segments I mean the cloud data backup and data archive grow growing more use cases but very new one showed up pretty recently and uh this year during open infra conference Nvidia actually released their statistics they use uh about one petabyte a day just for the a their AI training this is amount of data that they generate daily and well they have their own uh cloud storage uh that they use with open stack Swift and yeah this is why I think that this graph is actually uh growing exponentially I mean right now if you can see on like 2023 it's maybe like about onethird or maybe a quarter from what we'll have uh one quarter missing from what we have by 2025 but especially with AI I think that would actually probably grow a little bit further than that so yeah we need we need more storage takeaways from uh everything set up above uh cloud has been a force for centralization and oligopolies and we're in decentralized space we all know why centralization is bad especially on like tornado cash to be example here pretty much no matter how much people would tell the freedom of speech your data is private and secure if someone comes to your company with like us Court issued subpoena and tell us you need to delete this data or show it to us or give up your service to the government it must probably will do that I mean it's not only in the United States but I mean we have so much infrastructure in the United States so uh it is relevant and then yeah oligopoly being probably less concentrated problem in the central in in the centralized space because yeah the worst the oligopoly can do is shift up the prices but I mean there's so much so much shift you can do until you reach the house market pricing right so and the other thing is that open source actually helps but it is not enough because right now open source uh solution exists open stack and you can spin off your Cloud you can use it in your Cloud but still people do not want to run their own software Amazon is so easy and convenient why bother you just choose Amazon and there's not so much of competitive market it's just being centralized little by little every year so yeah um centralization boo oligopoly boo uh now is the time to build a storage and yeah here is probably another question to raise uh why there is a need for another decentralized storage because okay there there's obvious that there is a problem in the web two space but uh how much the problem there is in general and what I think that like very simple truth is if you run around hackathons or just o oversee like other projects uh in decentralized space it's very clear that a lot of people still use Amazon because of how convenient and reliable it is and I mean just clearly lack of an adoption of decentralized storages in decentralized space and yeah this is pretty much while we think there is an is still and why there's a new storage coming to town so this is uh a short list of our thinking how it would be the best to take something great of web 2 and web 3 and combine that so why everybody keeps using Amazon is because it has this four properties which is pretty much resilience availability capacity performance so this means the ability to write data when you want to write it to read data when you want to read it anytime then this service itself has to be available and you have to be able to put any amount of data that you need in there and obviously performance needs to be up toate so it doesn't take Infinity to load a small text file so and then comes web 3 and web 3 brings to the table immutability and verification of ownership and content just pretty much you can still have have private data but have a permissionless access in the network and yeah this uh decentralized projects also and web 3 brings a lot of stake into open source development in general because it's a it's a bad practice to have not a private repository in the decentralized space so that's also cool and yeah obviously web3 brought like exit to Da option currently so transparent Community governance and of acceptable and acceptable user policy is so just making mixing that all together should be great for the storage and that's what we're trying to do uh yeah let's have a little bit diving in what we're actually building the project is called on makina that's what I'm talking about here all this time that's what we're building yeah I'm a co-founder and the product manager of this and yeah we're building on top of near uh I guess would be nice to say we're kind of Layer Two to near uh we use it for authentication services and Ledger uh as of now near has this awesome account system that uh is kind of account obstruction where you can have unlimited keys underneath One account and every key has different access rights towards this main account well then also near uh super scalable in itself and the most importantly it is very cheap so that's why near uh on the left uh in storage itself storing data on chain is uh definitely not a good idea and uh that's why we're not doing it we have special dedicated storage noes uh and we only do Ledger things on near but everything else sort of chain and on the right right there there is reputation monitoring so this is something we called statistical reputation service basically like a validators who will oversee that nodes actually store what they say they store and uh they can track like how performant nodes are are they available at all and all the other things like that um did I miss anything nope all right let's zoom in to authentication for the node uh yeah pretty straightforward there is a user he logs in with an ear wallet very simple into our client app uh we check the credentials from near because it's also a key system we can just create a GV session talking for them and then they can use it in their console to send and Storage things to our storage service that's pretty simple for the node operators uh there is additional step that includes staking so it's pretty much the same you go to the client top uh you verify that you are who you are with near wallet and then you need to stake some tokens because well um obviously you need to put something to to make sure that people actually care about your data we need to take something from people and uh not give them away uh until they behave properly so so yep that's a short overview general for storage components this is the thing that was on the right so we have proxy nodes that whenever data comes in it goes to the proxy nodes proxy nodes uh make sure that replication replication works well and put it on the different storage nodes you can see these are all the nodes they are decentralized they can be in different charts they can be in different continents they run by different people and proxy just make sure that data is replicated well it puts it on different nodes in the different uh to the different people and in case node goes missing there is process of rebalancing just putting the pieces that that node owns into different other storage nodes that still persist all right okay now is the time to demo and I probably would need some time to Tinker with the screen so please bear with okay uh now I can share it my screen and you can follow up I hope where did okay here we are cool okay so we have two parts here first one okay this is the wrong tab first one would be the part on how you can okay there we are how you can actually spun off your own nodes so this is pretty simple uh you log in with operators on minina Doo it looks kind of like this we have a near wallet obviously we're on a test net uh you look in with your near wallet let's try my near wallet okay not yet we have ravioli. testnet which is my account you just log in Connect all right so this is the interface for the developers and uh if you're node you can just shift to operators on ma. and you will have all your nodes right in here you can add a new node let's say Oro demo yep create sub account this is what I've told that near already has this kind of abstra account abstraction thing because you can for one near account you can create uh pretty much infinite account amount of sub accounts so here we go it's created while it was doing I should actually have shown that you would actually stake some amount of of near into the contract to become a node if you don't have near that will fail so uh now you can launch your note and join the network for that we need credentials file here it is and uh we need to put in on our server where did that go oh this is so hard to navigate here okay so I have this server set up in this domain address uh and let's just go straight there all right we need to actually copy the file which I forgot so okay on Marina Json and we need to copy it to root it 145 40 105 83 okay hold done and is it here y Json right sweet so what are we doing next what is WR here we need to create a folder to for a dockin container to use it then we need to set access right for this port so the docker container can actually use it then yep we need to mount a Docker image which I clearly need to to copy from [Music] here oh yes this one as well we need to move the the file into the M folder and mount the doer image so the image is actually already there and the server itself uh this is just very freshly uh installed server with wuntu on there Ubuntu has pretty much nothing but the docker and that's a different different path for the image okay load in the image now uh in the in the admin page we can actually see there is a command pre-made for the dockin container to run which we can just copy paste it's pretty much set an environment variables and running the image under the on Ma name and let's do it this is where the magic happens it pretty much works by itself right now it it is about to sync all the other notes it will take a little bit of time to try and make it work a little bit faster we need to switch to uh the user side of it okay this is where this is the customer side where you can actually store the files you can authorize within the app it Al features the settings page where it's very easy to get your access keys from in Cas like you can also use use it console for it but it just nicer to show it from here uh what I do here is uh I just export my account and export uh the token which is nicely put in here for copy pasting and let's get to another tab of terminal and try to throw in some files in there so first I need to authorize and then I have a small script uh in here I don't know how good can you see that okay so in here what we do is I have a bunch of files like a thousand files and I just put them into uh on Marina and then we can see that you can actually see them in a different nodes Let's uh I don't know let's modify the file name so we can see it's actually working right now let's do file one numbers 3 four 5 six s okay and pushing sweet there's so much tabs right this one so you can see here the file are actually adding up and if you can see by the counter because there were files from like 60 to 69 not all the files end up being on this specific notes but uh I mean they show up it works yay all right um do I want to show something else yeah I also had another node here we can push more files and see that actually different noes would have different files here we go so here are some of the files that I just pushed recently so this is another IP address here and then one 18 one and this one a with A3 so we can just see that the difference file different files end up being on different notes it has a little bit more logs that I tried to push files earlier today and it was replicating it from the other servers you can see it here and back to this note if we just wait a little bit uh we can see it has it just continues to sync up the files that were previously pushed and yeah I guess this is our first client uh that you can use as a node operator we're very early in project so we're particularly happy that replication actually works all right that's how that's pretty much it with the demo uh where is my presentation nope this is my presentation slideshow loading yep so demo is done and yeah because we're so early early right now uh I would just want to outline some principles that we're building thinking of and maybe something like unique selling points at this moment so we're fully up Source obviously not in use for a decentralized project uh with us anyone can be an not operator obviously you'll have more profits if you're like a huge d Data Center and maybe it will not be very reasonable to run it on your laptop but still everyone could be a node operator and we just seen that it's pretty easy to run Docker uh on your laptop or from your laptop on the machine uh in the cloud Center then yeah predictable pricing which is always a problem with decentralized storages that uh they have their own token and then the pricing is in this token and then token shifts depending on I don't know news how successful the product is and so on so for now we're going without Tok that's why it's going to be very easy you have the demand you have the supply Supply is very reasonably priced and you always know how much it cost to run your data center so just as in web 2 predictable pricing yeah because we're a decentralized user data is actually private it can be tinkered with if you Tinker with any data just on your nodes uh then it will not be possible to uh redo the file that was because the file is uh using raser RoR cards the file is actually split into separate components if one of the pieces of file is tinkered then you cannot replicate the original file and the node gets slashed uh so that's that and full automatic replication obviously if the node is down it just automatically replicates all the data that was on the Node using the other pieces from the other nodes and yeah one simple API for devs to Target actually the thing that I didn't show up today but maybe you can check previous demo that we did at Denver we pretty much have um five commands that you can use which is to put files to get files and it's just that you create coner you put files and that it works and uh yeah that's that's it uh please check our website if you actually want to try and run your node en Joy test Network please reach to us directly uh for everyone else who wants to try and use it for their products uh use it through IPI just go to the website we have a open test uh the for for beta testing so you can apply there otherwise here's Twitter um and yeah now is the time for questions yep uh so why should I use the on Mah uh over RV Shadow Drive uh file coin CA uh and a lot of other coins like Storage Solutions uh well there are a couple selling poins there I would say probably were faster because where you put data you can get it pretty much instantly then uh we have predictable pricing so we you don't need to buy any tokens or shift any tokens or have like some intermediary in between so you can use their service you just use it I think Shadow Drive has predictable pricing and RV has the you know quick uh broadcasting uh I'm not that familiar with shedow drive but with RV uh I mean they're Perma storage so it's more expensive but you have this benefit of storing your data forever but then do you really want to store all the data that you put there forever I mean yeah it's different use cases but all right all right thank you for question any other yep can anyone not or only storage notes uh okay can anyone run the proxy note that's a good catch uh so as for now we're running the proxy node because we're early stage project but eventually yes uh for now we only have uh the software for for the node and the software to like use that node so just stay tuned tuned we'll get back to this yep I I can when I stored the data do I have a choice on how many notes it's going to be stored because you know do you know if if it is only on Two on Two notes whatever if one gets tinkered or both of them get get get tinkered then you know that the data is lost even if you replicate right sure sure uh actually the uh thing that I missed a little bit uh there will be a concept of shards as could be for example geolocation of where the shards exist because gdpr compliance is a thing and you can actually select the shart where your data is stored and uh close close close I need exit Escape in here no in here there's actually a concept of shards that you can check and you can just pick which one you want like globally you don't care or you want Europe or you want some specific thing and Well for now we're only thinking about the uh geographical points of it I guess there is a possibility of actually putting shards into some different segments of I don't know uh for example how performance you need your storage you use your storing the backups then you can do less performance for now no it's easier to build replication if it's like global if it's applies for everyone without choosing just oh that's why the statistical reputation service is there this is kind of validator that rechecks the data often and tries to replicate it if it doesn't replicate it shut down the note so there is there is a third party that checks that right more questions uh okay so this is the last one because we we are running of time so let's let's be quick so what prevents nodes from hosting the data on centralized storage Sol storage solution like AWS and then just faking it to the uh proxy notes or validators expenses okay all right quick answer thank you very much Paulina and uh let's give a clap for pina thanks for having [Applause] me
