# ETHWarsaw 2023: Paweł "Pepesza” Peregud + Michał Chiliński, Octant - Staking 100k ETH using QuebesOS

- Channel: [ETH Warsaw](https://streameth.org/eth-warsaw)
- Date: 2024-10-07
- Duration: 1:14:19
- Watch: https://streameth.org/watch/yt-tEWCblSHIms
- YouTube: https://www.youtube.com/watch?v=tEWCblSHIms

## Description

Workshop - A workshop by  Paweł "Pepesza” Peregud and Michał Chiliński from Octant. This workshop shows how to stake ETH using QuebesOS.

Follow us for more updates: https://twitter.com/ETHWarsaw

## Transcript

hello uh my name is uh Pavo perut also known as PESA and and I am mik or Mike chilinski together we will present you our uh the way how we sted 100,000 if um of Golem Foundation funds using cubos at least it will be our journey and path which we decided to take on that journey to make it happen and in fact we are exactly at the moment when all of those HS are starting to be active on the network so we are really excited right now to be able to speak about it exactly on that moment if the phone will ring and mck will disappear in a puff of smoke uh the servers are down right yes the truth is that when I feel an email is coming to my email that are probably notification from our Network Operating Center so it makes me really stressed so please forgive me yeah so first some introduction you've probably heard about the merge it happened on 15th of September last year and the two networks which previously uh was like independent from each other uh if1 and if2 got renamed into execution layer and consensus layer uh execution layers so GF and Friends they've dropped uh the fork Choice Rule and started using consensus layer for tips through the socket so uh now we all of us who are running a full Note have to run two clients together and from the 15th of September uh the new network is called uh just ethereum 2.to after the merch proof of stake Network yeah so uh some facts or maybe judgments about proof of work and proof of stake as you can see like this stuff is likely the most technical thing you will see in this presentation um that's not true in my part there will be a little bit of technical part so uh but like from our point of view uh if users the most important thing is now we finally can uh be block producers from home right we no longer need to have those uh pesky energy intense uh mining rigs uh anymore and uh I'm running uh staking from home it's like two small boards each consumes around 15 watts of energy and one is a backup for another and the up typing is really nice so it's it's a pleasure nowadays yeah so as I've mentioned this home staking is easy right like I can do it a lot of other people do it but uh there are some uh problematic things there uh first uh now so instead of basically running a really hot validator machine you have now a key pair which uh can uh sign can sign a transaction which will damage security of your funds so those keys are hot that's one problem you usually don't store your keys hot if uh any amount of money is attached to it and the other thing if you will sign at the same slot or same height the similar messages you will get slashed so this is like the main problem uh yeah but like uh ethereum designers they've made uh some they give us some hints of how we should run those networks like in you can see those hints in the rewards and penalties uh structure um there is a small reward for attestation it's like three steps ahead like two cents at the moment there is a really small penalty for missing an attestation uh you attest every um EPO not every EP right six minutes yes yeah six and a half minutes yeah uh and by the way to attest it you need to be online I believe for the whole duration of OC because you are a part of committee and you need to sign every block as far as I understand but like this is a detail right but if you miss this station you do it two steps back so it's like you can be profitable uh or at least uh net positive uh even if you are like half of the time offline but there is this huge penalty for double signing uh it's one if and it's ejection from the network plus some time in the queue so you can uh really feel the the pain there so so in fact this penalty for double signing uh it's not only one if because in fact you're losing all potential income and you are losing time which will spend in the queue trying to enter once again for example so it's much more than one if and additionally this reward for attestation it's only for attestation and signing attestations but there are still probability of signing a whole block and this income is like to magnitude of order higher yep but still like the uh message is pretty clear right you should care about uh not double signing and every other thing is kind of a bit less important yeah so uh how the network look at the looks at the moment uh we have a huge amount of active validators uh together they represent uh 25 almost 25 million a being staking right now some of them obviously this this existed amount is not a part of the active set and only like 300 of them got slashed so if you will like compute the fraction it's really small it's like uh 3/10 of a percentile so it's really really small fraction and uh while we don't know the story behind each of the slashings we know some of them and U those those are basically in two categories uh it's misconfiguration and it's software errors and as far as I know it's mostly it was misconfiguration so people were trying to fight for up time they uh basically set up few notes right and after that they've uh turned Keys active on two of them and they got slashed so usually human error as a most of the probably worst cases in our it sex scenarios and the weakest Point are users so still here misconfiguration is usually users fault yep um yeah so message is clear right so uh Foundation stes 100,000 if um this amount uh actually enables us to make it properly to stay properly without taking uh shortcuts or like reducing the security of the network and on other hand like we can do it uh still like as a somewhat private businesses in we don't need to uh worry about the security of the ethereum network because this amount is like is is is nowhere compared to like uh 25 million if which are currently in the active Set uh yeah so what we what what ways we had like what what what were we considering when we were um uh preparing the staking like those four me methods uh they um are on a spectrum right so uh staking as a service you don't run anything you just give your if to some uh organization or possibly you give them your valid validator Keys uh and then run they run everything for you um liquid staking as a liquidity provider you just put your money into some contract and then uh someone else will will run valid for you so it's a protoc call it's already like a a bit more secure then there is solo staking and there is this liquid staking as a node operator for instance we could be as a foundation would could be a rocket pool operator uh this would make us a bit more money obviously um yeah but in general all of those uh apart from solo saking they come with added complexity and complexity uh means more potential errors yeah so we went with solo staking low technological complexity it's really well aligned with ethereum uh design goals because like literally people working on ethereum Design This solo staking process yeah and uh we had some uh experience in the team doing that so this is what we went with um yeah so a bit more about staking uh as usually it goes so you have your full note uh you create validator Keys you deposit F you wait in the queue and your validator are active so really simple right the only uh tricky part at the moment is this waiting in the cube because uh it takes like 30 days um but still like you need to uh pay attention a lot of attention so what I could say about this waiting in que uh is the reason why it's really hard to validate some ideas and make that quickly because uh if you would like to put some validators online it's not that we put them online then we are able to unstake them quit really quickly and then return very quickly it's one month the waiting in the queue uh so from the point of view of the project it was important to start staking as fast as possible but then it doesn't and it's time for us to verify on living organism some ideas so of course we made test on small amount at the beginning but then there was this waiting time really hot time when we tried to prepare everything for uh much more viid dator than before uh to see how it will works because of course some people said we have test networks and we could try on test net not the main net but then there is a question are we able to try on the test net 100 uh tokens uh 100,000 tokens from that Network that's not true it's not possible to check the bandwidth and how it works on the test net so our test is make on Main net of course course we made different steps of that but this waiting time in Q was quite crucial uh part because when we already were in the queue it means that it have to start working fine already after we'll be active because then any change in the whole setup it will need exiting or maybe just waiting and missing attestations knowing that in the final uh setup it's probably more important to be protected from slashing not from this downtime which could of course at the beginning decrease your amount and give some net loss yeah um so uh if we would like go a very Nave way uh this how our components would like uh would look like so basically you have your um GF it's execution layer client you have nimus it's um consensus layer client they communicate through a socket so basically nimus leads G and tells them what's the is actual head of the chain G feedback information to nimus telling nimus about the um valid um validator set and there is this small uh process which is called validator client which is spawn by nimol and it actually uh stores Keys your validator keys so uh nimol already does some effort to secure validator keys by spawning a separate um operations operating system process which is like means separate uh memory stack and everything uh but as you can see like if someone will uh break into the gaav or nimos um they basically overtook everything because like priv escalation on in the Linux is kind of easy and it like those exploits are quite cheap on the market uh yeah so the thing is uh we can actually take this a bit of further in terms of like securing the keys yeah uh and here this this cubes Parts comes in exactly so I will ask the question uh who out of our great Auditorium knows anything about cubes who heard about cubes who maybe is a user cubes okay that's great at least half of you uh of course at the beginning I will tell you what is cubes and then I will try to explain why we decide it to it means why they decided to use cubes for that because I'm from itl team and we are developing cubes and it's already our 11th anniversary that year since the first release of the cubes so it's already quite mature project and of course we are usually a system for uh desktops or laptops but I will try to show you how cubes in that scenario uh is uh quite useful not only because of security which is uh buil-in Cube cubes uh additionally it's quite nice because we could use features of Cubes to make the whole staking easier and uh more secure from the point of view of for example probability of slashing probably we need to clone the screen and that will be fine so I will show you at the beginning couple of slides from our cubes presentation and then I will show you exactly how it looks on our infrastructure right now and on the final step of our Workshop I will try to show you uh our uh real dashboard with Statistics from one of the machines which is working right now and I hope as usual live demos are not working on the conferences but I hope that this one will work uh okay I don't know why it's so small but in general cuos are reasonably secure operating system so of course when we are talking about cubes we are saying about the state of the world and the threats and to be honest cubes is the solution made on compartments what does it mean it means that we know that it's impossible to create perfect software software which will protect you from all potential threats so in fact we know that everything could be exploited even small weakness in one part of the system it could be driver in your network card it could be some flow in the client of ethereum or it could be some flow in the parts of the operating system for example open SSH or whatever if it's one operating system one M analytic system then andhole will let you in so in fact in the best scenario attacker could for example shut down your machine so then you are suffering from missing of attestations from downtime but in the worst case attacker could steal your keys and could start double signing so could start slashing your tokens of course if you will realize that something is slashed then those validators automatically are through out of the network so but if you have for example 1,000 validators or 3,000 validators then it means that you could lose because of slashing for example 1,000 if it will be quite pity to lose so much so of course uh what is the idea this is the single system uh in the cubes we are doing that of course humans uh we are doing that in the idea of compartmentalization it means that we are trying to divide the system in separate parts we are putting in one part Network Internet controller Nic and drivers to that in different part we put for example different clients and I will show you exactly what is the scheme for our teching right now and cubes is based on is based on virtualization so it means that we have hypervisor which is working uh below as the bbon of this system and on top of hypervisor we create services and we made all of this bunch of VMS working together usable as a single operating system so right now for basis for the cubes is Zen hypervisor it could be something else but uh since this 10 years we are using Zen and we made a lot of services which are connecting and communicating inside with very defined very small communication layer it means that there are so simple protocols for communicating them that it's really easy to audit them and of course users of our solution there are different companies I will show you some uh logos and some project they were audited and for those Solutions we know that okay this layer of communication uh is really fine audited and we could trust that and of course uh we decided to change this tripical system uh to the cubos way where we could create not only small compartments which when they are exploited are not so easy to affect the whole system but additionally we are able to create policies allowing communication or not allowing or changing the way how we could communicate how we could Monitor and verify what's happening so in fact we are everything uh dividing but why it's Unique it's because we are able to configure everything we have explicit policies and then we are able to separate configuration from the data and how it works and of course origin of cubes are the systems where we are sitting in front of the computer it means that laptop workstation or desktop but we are able to use the same scenario on servers so of course it's still possible and uh why we decided to use cubes of that because at first there is this security of operating system that's obvious but additionally what are the advantages of using cubes like solution for staking is that if you are going to try with the high availability scenarios you have additional features for protecting of double signing because of course at one of the level when you have a lot of validators and that's the case which we have right now you start to ask what is the cost of one hour of downtime or one day of downtime or one minute and then you're able to realize that this cost is in fact counted in thousands of Euro so these thousands of Euro could be used in Octan project for fueling development of the society world society so it will be waste of resources to not gaining that profit so of course we started asking okay what we could do to make the better availability and higher up time then we should find the ways how to for for example be sure that when we are moving keys from one machine to the different one that for sure on the other one there won't be any copy of those and for example with cubes we are able to not only turn off part of the system because in fact it's everything enclosed in one VM we are able to remove AET controller out of that we are able to destroy the virtual machine so in such way we are sure that there is no more copy of keys and no more this part of the system which was resp responsible for signing Parts if you have monolitic system where you have the whole node working at once it's not so easy to destroy those in our case we are able to just uh remove a ternet we are able to remove the whole part or we are able to restore the new signing component because we have uh special templates and we could just copy template of course empty template and then put inside keys and start signing for example different validators or restore those so thanks to cubes we have this additional features for for example updating parts of the system thanks to templates additionally we are able to create secure scenario for copying Keys between VMS and in fact different machines with the idea of the CER protocol and finally of course we are able to uh be sure that the validating client it's not working anymore when we don't want to uh in general Cub is quite simple for management is because we Bas on policies uh user interface is not visible when we are working on servers but still if you like you I you able to log in to those uh and now I will show you uh how looks our setup with uh this exact scenario but who is using cubes so if you are not familiar with cubes maybe some quotes for example Edward Snowden uh he is one of the users of Cubes regardless if he is russan spy or not I don't know if you believe that theory but in general uh he said that he gives him power and that's true when I started using cubes I right now I feel uh strange when I need to do something uh which is quite related with privacy on computer which is not working cubes for example right now I'm not able to log into any secure service like banking or something else on a different computer it means that I feel like in a car without safety belts you feel somehow naked and it's the same feeling which I have right now when I'm not using cubes so of course I sometimes I am not able to log in somewhere because I don't have my cubes and you see different quotes maybe what is quite important there is to mention uh one person from the community if Community who said something about cubes there are different companies uh and Foundations which supported cubes which are using cubes and uh fueling with donations uh of course cubes so let me show you right now okay some quotes Okay so maybe this one uh I don't know vitalic bin this guy said once trying out C and that's is distro designed around increased security and surprisingly good user friendliness it's something what we got as a sometimes comments that cubes is really hard and how should I use cubes so in fact he said that it's fun enough and for Server operators of course it's really easy but that's true that Cubs recently is used mostly by expert users maybe not expert users more power users let's call them it means that users who are not afraid of Linux uh so in general when we are thinking about Serv servers it doesn't matter because your s admins probably still are good enough to operate with that you are not afraid of command line interface because Gaff nimus never mind whatever usually you just Spa them with the command line and additionally there was uh one other quote from the same person uh he said that if someone has one gig Euro or dollars they could p that into cstyle secur operating system and it will be something like Manhattan project for cyber security so it will be nice feature I believe that there will be someone who will pour money into the project but let's go back and see how we are doing this solo staking right now okay so here we have two general pictures of cubes in general this one example of architecture of Cubes it means how we divide different parts of the system based on the level of trust so we have level trust which is safe and ultimate trusted it means that we don't like that level in The best scenario there won't be any part which we need to trust but of course there is a hardware there is user and operator and of course there are Parts which are unsafe and untrusted so in general as you see we need to somehow trust into our hardware and that true right now we have uh our server room and uh we need to trust operator adap server room and the hardware which we bought for that and of course there's different ways how to make sure that your Hardware is secure that you're not buying them directly as a foundation advertising that we're going to stake on that Hardware something uh but of course it's the first step with this Hardware related paranoia that you just taking look how you obtain the hardware uh this plot is in fact on our website cubes OS cubes dorg and then there is this additional layer of course here this H hypervisor we have Administration VM and one level deeper we have different parts and for example as you see here we have Network VM and what is important in our scenario Hardware like network interface it's not directly connected to the application VM here this application VM here of course you can see some clients uh applications but in fact in our scenario right now here we have G as a additional VM it could be Nimbus it could be validator so of course we don't trust our Hardware our ethernet connection and ethernet card so it EET card is connected to the cisnet VM and it's our first step of uh security and our first step of Defense it's this sisnet so we have physical hardware and Driver moved from the Dom zero so it means this Hardware layer directly to the separate entity and then if someone will be able to penetrate this net then there is the second level uh of protection and then we have CIS firewall is a separate entity separate virtual machine so in fact here you have Hardware Hardware is connected to the cisnet here is the internet connection initiated then there is the firewall as a separate virtual machine and then firewall is the place where all of the applications VM are connecting so for the very simple scenario when where everything will be working in just one virtual machine like here we need at least three virtual machines for that so Network stack firewall rules and finally applications but of course we are going a little bit further and what's is important when we are thinking about cubes and probably part which is the trickiest for users is that you have to configure this part when you put your own services so of course when you install cubes by default you get something like that so it means we have this setup set as a default if you're are not pro expert and if you are not uh checking the Mark I'm expert please do not configure default VMS for me you will receive something like that and addition alties which are quite characteristic to cubes are disposible VMS it means that if there are things which we would like to try we are able to just boot up one virtual machine to make some oper ation and then destroy this virtual machine and of course for easiness of management there are templates so those templates are prepared preconfigured pre-installed with some software virtual machines which then are started in this app VMS the greatest pain for users sometimes are okay how I am able to divide my scenario of usage into different VMS so of course when we are thinking about the user in front of computer computer we could divide it into the work with different areas related to work then we could divide it to for example personal staff or some red browsing of websites with funny pictures uh everything could be divided in terms of solo staking it was much easier because in fact we have just one meta service staking but it's divided into three parts so let's take look what are those parts and here is our I hope it's visible okay so in fact it's our scheme for solo staking o cubes what it means it means that here is our machine this blue square is the concept of our hardware and then inside Hardware of course we have cisnet and here we skip this firewall partk but in fact of course there is this firewall somewhere so there's the Cy net and firewall net and and here's the part related with uh staking and here we have three separate virtual Machines of course one virtual machine uh is execution client and here right now it's GFF but of course it could be something else with cubes we are able to create template and next virtual machine with uh for example never mind and put it as a separate entity there so for example switching to the different client will be quite easy then as a separate one of course there is consensus layer and right now it's nimus and part of the Nimbus it means that separate nimus validator client with keys is in a different vm2 and what's important of course consensus layer and execution layer those two need to be connected to the internet so uh if there is internet then there is connection to the cisnet so of course as you see there are already arrows connecting G and nimus with internet and there of course we have some bandwidth for one machine it's something like 60 megabits per second so to be able to have more machines and I will tell you something about the numbers uh in a couple of minutes uh it means that this connection need to be stable and you need quite high bandwidth there but in term of validator is the hottest point from the person of security there are the keys private keys for signing so it's the part which you would like to protect and in our scenario this part is not connected to the internet at all not to the network it means that as you see there is no connection between cisnet and EHD valid because we are using our internal mechanism for communicating between validator client and consensus client we use qer exec qer exec is internal way of communicating between machines in cubes and we are able to put into CER exec TCP connections so in fact there is the special layer of communication on cubes which use for that and of course we are able to Define policies on that and even we are able for example to create our own wrapper for those connections and we are able for example to check what is put through that pipes so for example if there will be need for that we are able to create our small proxy which will be reading that traffic between signing validation client and consensus client and check if there won't be any requests for the same for example EPO slot to be signed twice we are able but right now we decided that there is no need for that because in nimus there already some features which prevents that solution but what gives us cubes cubes gives us the POS possibility to connect those two boxes through our internal mechanis which means that we don't have to to use the network stack of the operating system uh which is the main operating system of the machine so for example in the scenario of high availability we lose connection to the internet from this machine and this machine is looking that okay I don't have access to the internet for longer than some thresholds like 10 minutes 1 hour or 5 hours and in such case if this threshold will be reached then the machine could for example delete the whole VM delete or turn it off there are different scenarios for that if you would like to be really really serious about security and we need to be sure for 101% then we could delete that machine if we believe that nothing will turn on that machine we could just uh turn it off but when the machine is deleted or turn off it means that we could boot up or already put those keys into a different machine in different ht8 node which was working already somewhere in the network waiting for this disastrous scenario when one of the node was lost on our site and what is good about this scenario there are still working two clients so it means that for example if it was only problem with some tractor destroying fiber optics somewhere in the ground and then the internet connection will be restored we have already everything almost configured there it means that there is already one terabyte of data we've there's already couple maybe 100 of gigabytes of consensus client data so we don't need to synchronize once again since the beginning this machine is already almost ready for operation so then we are able to for example put here different amount of validators validator clients and we are already working it means that thanks to this mechanism which is in fact thanks to the cubes we are able to quite easily create high avity scenario which will be still on one level more protected from slashing than the normal uh setups where everything is in one machine and of course we need something for monitoring that stuff because it's not really easy to put everything in the computer turn it on and believe that everything is working at the beginning three months ago we thought that it will be so easy but unfortunately all of those clients they are not so well tested especially not well tested with so many validators because in our case we are talking about over 3,000 validators so some of the scenarios or maybe amounts of the RAM memory consumed by the processes were unknown to us especially because test net was not able to provide us any real life scenario it means that number of peers in the test net and uh volumes of traffic they are completely non-comparable with the amount of traffic which you have when you are staking thousand 100 thousand of if so for monitoring that we created separate cisnet management part and this part is connected not only to the pp uh P2P network but is connected to the VPN network because this part this cisnet is able to connect with the world with the public IP address but here we are going and we are allowing only P2P traffic from internet Network so we are able to verify what is the traffic and allow only that traffic but if you would like to ssh in the computer to make any management if you would like to see what's happening inside we have some monitoring so we created additional monitoring VM here is our graph on now working with plots and graphs and of course to access that we have VPN network so behind the VPN we are able to SS into the management VM uh into the whole cubes and operate it and of course we're able to see how it's working and on the next step we are going to create one master dashboard for all of the machines because of course we decided that it's not the best to solo stake on only one machine we decided to use thousands of those so of course right now this is setup of only one machine so right now if there will be for example 10 nodes and 10 machines then there will be 10 monitoring uh VMS and 10 dashboards so of course if you have a large enough display or couple of displays then it's possible to have 10 plots but of course it's better when you are able to merge the most important metrics in just one so it's our next step uh which we are going to develop for that reason but in general what were the main problems which we faced when we are when we were preparing everything when you are dividing Hardware into different VMS you need to think about physical stuff below this physical stuff below is something like storage what should be the amount of storage which we assign to the G to Nimbus and to the other part of the systems because we need to take care of that because of course the blockchain is growing the amount of data which you need to store is growing with cubes we are able to quite easily divide those parts but at the beginning we need to prepare that additionally there's there are of course cores CPU core for signing and for calculating cryptograms and there is R andom ra Ram So Random Access Memory and for example G is a beast which is able to consume all of the ram you have on your system and additionally it's able to consume all of the swap which you have on your system um but what is really surprising when we have when we had setup with 128 GB of memory dedicated to G and one gigabyte of swap we found that g consumed all of the RAM and consumed all of the Swap and then stopped working so we increased amount of swap to 10 gbt and we restarted everything and suddenly G consumed 128 GB of RAM and then nothing out of slap and we made this experiment couple of times and always when swap was really small G was a B to consume everything uh so finally we found that it's enough to start G with only 32 gbes and it's working it means that most of the ram is consumed by cach and we don't know why when we have uh our SSD drives which are so fast that they're almost the same speeds like Ram when we made synthetic tests on ram ram disk and ssds there are something like only 10% or 15% of difference so when you have SSD drives which are able of writing 7 GB per second it's really a lot so probably uh of course network is not giving you more data than that additionally with G uh there is this small problem that we don't know exactly when G is making snapshot of the files on the hard drive so sometimes when we were killing G to check what will happen if there will be loss of power supply G after reboot needed only 15 minutes or half an hour to synchronize with the network but sometimes we got message that there is corruption of data for free last three days and we had to wait a couple of hours to restore that so of course uh is the great to have multiple of machines to be able to move validators in such situation when you are when we are restoring them uh and in terms of for example Nimbus and what here could be a tricky part when we are connecting think with validator we use sockets from The Cure exec and sometimes nimus is spawning so many sockets that kernel is not able to properly close them and is closing them um in the smaller Tempo than opening the new one so you are able to consume all the Ram or you are able to find the soft lock on the Kernel and then just lose VM with consens client but luckily our team was able to find how to patch Zen and how to patch kernel and we did it and it looks that it's working fine right now so our first test what will happen if we will spin up 400 validators on one machine and it was able to work for 4 hours after 4 hours we consumed all of the slots and Cal stopped working so that was fine that was planned that it will stop working somehow it was after 4 hours then we put our patched kernel and right now it's working already for one day and a half because in fact it's everything really hot and new today uh during the night probably 1,000 validators more will be online and of course there will be uh more interesting to us it means that on more graas there will be something interesting to watch and this is probably the moment when I will show you how looks our grafana because there's already automatic suspense yes thank you uh and then there will be time for you to ask questions so right now this will be this tricky part when I will try to take my laptop connect to our secure VPN and then give you idea how it will look like and what is the final message that the H work which we are which we did and which we are still doing we plan to create a wide book about it we plan to create some how to tutorials manuals for the community how you are able to use cubes for solo staking and what we have learned from all this great journey of solo staking it means that we are able to tell you what could go wrong or we are able to tell you what could happen with different drivers of uh different different network controllers or for example how to maintain sync of clock inside the vmss because as you see in this validator VM we don't have Network so it means that we are not able to use Network time protocol inside it but luckily there is cubes and in cubes we have cubes sync service which is synchronizing clocks of the VMS in the real time so we are able of course to maintain the synchronization uh I need to connect somewhere if you could help me it will be great and uh of course we tried what will happen if there will be no Cube sync time sync service we found that quite fast VM is starting to have synchronization worse than 800 milliseconds or worse than 1 second and of course is the moment when consensus client is starting to shouting that oh no your clock is destroyed is I'm completely out of sync uh so then it was for us obvious okay so we need to start synchronizing the clock and it started working once again okay so now I will connect to our secure VPN and I will try to show you how it's looking okay I'm connected and I'm refreshing perfect it looks that it's working so right now let's believe that my cubes based laptop is able to show anything on HDMI output uh okay something happening perfect it's cubes it's working okay as you see it's a dashboard from grafana nothing fancy but what we are able to see here for example we are able to see this part of the test related of the old kernel on Nimbus here as you see is Network transmission rat so we are consuming something around let's say below 100 megabits per second per every node so is the moment when we finished waiting in our queue for over one month so is the moment when finally our validators started validating and then you see that of course the network rate started to increase quite quickly so we reach this plateau of consumption which here is uh around let's say 70 megabits per second and of course Nimbus started to create a lot of sockets and after a couple of hours everything crashed it means only one machine crashed and it's thanks to cubes that our G was still working there so we are able not to lose the sink of the GFF but of course we lose nus for a while then we changed the kernel we started once again the machine and then we had to wait for sink of the nimus so of course we decided to wait this time to check how long it will be although we rebooted the machine in something like 5 minutes it means that we were aware that probably something will be wrong so we were waiting we saw okay it's soft lock then we killed the machine rebooted with the new kernel and then started waiting for resing synchronization it went really fine after a couple of hours like two hours later we started working once again and here you see attestation rate it means that is attestation rate of the uh node which is operating here so we are somewhere below 4,000 attestations uh per hour so it means that in fact we are ATT estaing here with 400 validators and uh is it large amount or not we could take a look on the consumption on our machines and here are three VMS of course execution client consensus client and validation client so as you could see in terms of CPU com consumption we have here for example 16 course dedicated to that and there is a lot of space left so probably we are able to validate not only 400 but maybe 1,000 maybe we'll check that on sometime but right now in terms of CPU it's fine this machine here it's 24 24 cores 48 threads but from the security point of view we are using only course physical course we are not we are not using threets then memory usage as you could see this one is our G so 32 GB and of course more of the most of that is Cash uh and sometimes it's growing for example app is already something like 14 gabes so it's growing too but then there is a swap usage and right now everything is free that's fine when there is some problem or maybe when there's high peak of data you see that IO weight is higher on CPU it means that we're waiting for I our drives so these piak are related with the events like resynchronization or like here you see it was this downtime on Nimbus so afterwards there's high peak because nimus started to put data to our execution client and when we see consumption of memory on other Machines of course uh we see that it's really manageable in terms of our validation client is almost nothing in terms of uh consensus client here we put 16 GB but still at the moment app is using something like five 6 GB and when you see that there are some drops of RAM memory in the machine is because in cubes we have this uh Dynamic run Ram assignment possibility so we are able to change the pool uh dynamically between the machines of course we could switch it off uh for example for GE it's stable because GE is stable to consume everything but for Nimbus and Nimbus validation client and we have this active and here it's Network transmission so we see that for GF it's something like single megabytes per second but for Nimbus it's 60 70 megabytes per second and of course we don't have network controller in our validation client so here is no data so we don't have it and of course we are monitoring different parts like storage because of course storage is quite important when you have this whole block so as you see the right now we're able to manage that and thanks to cubes we quite it's quite easy to change the pools in different VMS so right now every machine has its own dashboard and our next step will be to create single dashboard for ofds and our next step is the idea how we are able to validate and check what is the real performance of every single validator what is the tricky part right now because when you look on the external Services Services designed to monitor your validators you see that for the basic free version you are able to monitor for example 100 of validators but there is a special tier for whales and with whale tier you are able to monitor 10080 validators so in our case we need to monitor 3,200 validators and of course there is no tier for us and buying uh wh tier times 10 times it doesn't have sense because in fact it still means that we have to look on 10 dashboards so probably we'll need to design our own solution with for example nbus node dedicated to the monitoring of the per performance of the validators not only rough estimation that what are the number of attestations per hour but for example how many blocks were caught or for example what is the real income because then of course it will be used by octant project at the end so we start first 400 validators is already working as you see it's working in the real time time offset between VMS 63 milliseconds so that's fine and during that night 1,000 more validators will be online so additional dashboards will be more interesting too uh and of course the whole project is going should we have some questions uh because we have like 10 minutes more sure uh let's use the mic because because we are being recorded and this is the only source of the sound uh so from what I understood uh if you go to the back to the dashboard uh so uh all of those VMS they are running on the single machine yes that's true so uh the consump memory consumption and also like CPU it's so low for the consensus layer that you are able to spawn so many VMS yes all right it means that right now uh I will return to this uh first dashboard because I tried to open for you uh the dashboard from the previous machine because right now we have two machines which are already working with the clients which already went into the online State on first one we have 62 of validators uh so it's the one which we tested everything uh we started with 32 validators it means 64 then we divide it into two machines we tested different scenarios and and then we started with this 400 and of course during that test we decided which will be the best hardware so when I tried on the internet to ask hey guys what Hardware do you need to put 3,000 validators people were what 3,000 we don't know and so we made some tests synthetic tests based on the uh information which we found and on those what we tried to Benchmark there and we realized that probably 64 uh gab of ram will be enough and 24 cores will be enough and that's true this machine with 400 validators it's working with 24 cores which are uh on and 64 ram but of course it's server machine so we are able to put 24 additional cores just populate the second CPU slot and we're able to put one tab of ram if there will be need but 64 is enough for up to 500 validators we checked that and it's working uh maybe in month for two we'll be able to tell you if it's enough for 1,000 maybe it's true too nice nice that's surprising low so uh another question in your opinion what are the pros and cons of this approach compared to something like um deploying each validator node separately for example in a cloud as in like a do Docker containers or something like that CU that seems like I think more common approach I guess I think that it's the main question meta question related are we going to trust to someone else and give them our keys or not so I believe that Foundation here decided that it will be probably the best to have our own keys with us together and to be responsible for what's happening there because when you have Cloud Solutions it's possible that uh when you are working on cloud there in fact there will be two copies of your machine because you don't know what's happening below the hood behind the hood of the for example Google Cloud it means that you see in your dashboard there's only one machine but what will happen below you don't know exactly what will happen of course one way is that there could be some malicious operator who could by intention destroy something or sometimes the backup mechanism could make the situation when there will be the ghost machine of yours which could try to connect to the internet and I know the of course not the taking scenario but the scenario of uh scientific data collecting service when of when one of the machines had two copies working at the same time because it was copied to Second Data Center and those mechanics made that the First Data Center not Switched Off the machine so of course in this scenario we the clients probably the second client with the configuration with the same IP address we'll have the problems with uh getting enough P2P nodes to start propagating but we check that even behind nut when we put one of our notes behind the nut number of pi were in enough to sign attestations so it means that even if we'll change IP address public IP address so this this Cloud scenario this machine was still able to aestate so double signing scenario is possible there thanks thanks that that makes total sense so uh what are your plans if you eventually run out of some resources and you need to spin up another machine uh how would you approach it in such in general our idea is quite simple for example right now we have uh X of the machines plus additional couple of machines which are our spare machines and we use them for tests like I said never mind test will be probably one of the next steps which we'll try and of course we when we'll see that we are going out of resources our next step is to power on next machines and additional thanks to our spare machines we are able to just increase number of Hardware it means that okay there's spare machine we see that 64 gigs is not enough so we are putting into the spare 256 GB we put additional CPU uh or we put fiber optic uh internet controller there to have 10 gbits per second we put that inside and we are able to move validators right now we built into the cubes a mechanism for moving Keys between machines because it's quite tricky too because right now when we are creating keys for the validators those are created on the offline machine like on the secure terminal so it's Cub's machine which is unconnected from everything and then inside the validator client VM we create DPG pair of keys which is our Transit keys so then we send this public key with Cub's internal mechanism which is uh unrelated with the network then we pass this public key to this our secure terminal we sign this transport package with the validator keys we put them in the validation machine and of course for the high availability we need to make scenarios when one set of validators is signed by two or at least three keys because we have spare machines because it's simple to say that when one of the machine goes down you need to restore validator keys for that if the machine is still working but you lose for example G or nimus you are able to copy directly uh with this encryption mechanis if not because the the machine is doomed and exploded then you need to restore those Auto encrypted copy so right now we have uh at least one per for every key validator key so it means that when we lost uh for example machine A1 then B1 or C1 from different side will take care of that and of course the whole setup is made in many sides so right now we have couple of service inite a couple of Serv site B couple of server side C and we are able to and in our future idea is that we'll be able to manage the load for example and the backup between all of these sides and of course we could put different sides under different jurisdiction in different countries to be able to maintain the operation of the nodes regardless of what happens around what about another mind are we going to use it in the near future answer is yes we will uh just for the guys yes it means that at the beginning there was someone told us it was the guy named PESA he said on my setup at home I is G so let's try with that and we start with G but then the guy named PESA said what about never mind maybe you should try with never mind but then there was a really serious guy called julan he said okay but there is no time for that let's try with G because it's working so we started with GFF and of course then the guy called PESA jumped in and said maybe let's try once again we never mind so of course with spare machines we'll be able to try that and of course we'll be able to make side by side tests and for example to prove that what is the amount of ram consumed by different uh clients and we will be able to ask you how to make the never mind better solution for that and of course it's possible to switch between different clients because in cubes it's really easy to create a second uh uh template with different client of course at the end there will be question if there is enough of memory there to be able to store on one machine different copies of uh chain but of course we are able to put put additional hard drives there uh and of course then it will be doable so with Cubs we are able to quite easy test different scenarios I'm still trying to connect to the old grafana but I need to take look for the very secure password to open that and I will show you a couple of days of grafana uh logs um okay what's funny that uh it's really interesting for our team too so it means that even uh mik who is the lead developer of Cubes uh is uh really happy or maybe motivated to work on that and uh he was working uh together with this team responsible for deploying that very closely and uh thanks to him we will be able to for example deploy quite easily patches for the Kel so it was adventure for the whole team okay so I think that it will be here ah yes I did it great so here are deploying this our older grafana dashboard it means that we are improving even how looks our dashboard between grafas oh no it's already updated great so team is still working uh so here you see this the different client it's our first client and this one is working with only 64 validators right now uh but the consumption of resources is quite similar um of course this memory is a little bit trick here because it it counts cash as a consumed memory so in fact it's not always true here you don't see any hole because this one was not affected for example we still try to understand what happens with our testation accuracy which is usually related not exactly uh with our performance but the performance of the internet network uh if Network and for example here atation rate you see those uh PS and Maxima and minimas are not related with the bad performance but here the fluctuation is really small attestation rate per hour is almost constant around 578 attestations per hour because here we have 62 validators I will try to show you how it looks like one one week uh I hope that my browser won't explode at that moment because I don't have a lot of ram in my laptop but for example here you see the moment when we changed n amount of validators on that particular machine here as you see there were 32 validators and then we moved additional 30 validators to the machine so of course we see quite large increase of network transmission there and uh on all of the interfaces of course number of peers were stable there but uh when we take look on the resources for example here okay here's huge peak in CPU consumption but what's surprising when we put 400 validators there was almost no increase in CPU consumption in the comparison to the 62 validators 62 because in fact uh uh we started with 64 then we made controlled exit of two of those to verify that we were able to exit and then we returned and even you see how stable is the consumption of RAM for the one week for nimus for GF and for uh validation client so everything looks fine and of course this increase of network traffic because of validators which move to that machine any more questions thank you very much thank you [Applause] guys that's true thinking about ster who is not so technical as for instance v m it was easy for but for some to be honest as I said it's probably it will be our plan it means that for me as a part of the Cub team it's something quite obvious that we should create white paper about it or the white book which will exactly explain which parts or even the salt formulas for that because everything was built in salt so you are able to quite easy to copy that and I think that's work we should be are from one inst to be honest uh if you're able to deploy cubes it means just install plain cubes on your computer uh we are almost there it means right now there is software which is putting commments inside your command line and doing everything with one click so right now when I will call M I need one more machine and here is entrance to the management interface of that machine and here is the public IP just power it on so for him it will be just putting the First Command pressing enter and waiting and that's all and this machine will download everything what is needed then the machine will start synchronizing and finally it will start validating so we are really close maybe there are some tricks which are important to remember for example Hardware compatibility uh it won't there could be possibility that for example your network card if it will be Co strange card won't be supported by then uh but in case of servers most of the cards are supported but for example it will be um Intel card uh I think that all of those works we had some problems with some broadcom chips uh but it was one of the potential problems but inine yes just Hardware problems but in general installation of everything after the whole route which we um go through it means that putting all of these formulas together and there are of course parts which you need to configure like putting these secret codes inside or generation of the key but if you are aware how the clients are working and what you need to Define which variables you need to fill then the setup of Cubes shouldn't be quite hard we are doing that in the one uh one click and just wait for the installation and I believe that finally there will be some octant related documentation for that uh um we just need for the support for someone who would like to pay for the food for developers to write that but in general I think that we are not far away from that coin maybe not cubes coin I think that some fat will be better because we are paying for food with fat cold hard cash but in general it means that as a cubes for many many years we are not really happy with crypto world with proof of work scenario and we would like to be really fair against the world so I think that staking idea it's idea which is much more aligned with the our principles uh in itl and Kos so we would like to be part of that change if it's taking and of course if there will be people willing to give us tokens which we could then switch to the money which will give our developers possibility of living uh then that's fine too so maybe we will be on octant too if they will invite us there have you got any money from a government recently from government not you're AG green it means that uh not cubes directly not we have some you founded grant for something but it's not from government directly anything else any more questions and it's still working there if not we can wrap up oh yeah sure and uh could you repeat that uh about e memory do you did you have problems with uh gav memory usage yes it means that uh G was consuming a lot of memories a lot of memory and in fact what was the most tricky part of g i not maybe I don't couldn't find that how to force G to make snapshot of the data database State because sometimes G after closing it uh in the unclean way needs only couple of minutes to synchronize but sometimes G is telling that one or two days or three days of data was lost do you mean synchronize after restart or yes synchronize after uh forced restart not the clean shutdown but after the forth restart and sometimes I see counter uh in the logs from the G like creating snapshot and I believe that this snapshot is the way of writing down to the files the state of the chain but we don't know how to force it to make it for example once every eight hours or or every day because when we tried sometimes this unclean shutdown was because of some problems with software or with the kernel or with the power supply but sometimes uh it was just we decided to put system shutdown service shutdown the Linux and system D looked okay I'm waiting already 2 minutes it means that I need to make force shutdown so the system D killed the process and then couple minutes later we just make start and the g told us okay I need three days to synchronize and we were waiting eight hours in system D there is some I don't remember exactly what is the flag but there is some flag configuration for graay fully closing and it's the key for get to add around uh 30 seconds uh to give that time to close the wall G gracefully in our case we give it 30 minutes it means that to be sure so even 30 minutes we put it for the graceful shutdown yes it means that uh once we waited really long uh and finally after something like 12 minutes it was able to close gracefully so it means that right now when we are doing our Administration task which are planned we decided okay it could be something like half an hour we could wait even half an hour because when we scheduled our maintenance for like 24 UTC then we could just start switching off one uh hour before or half an hour just to be sure not to lose the thing so much I don't think if there's need for Windows version it means that I don't know anybody who is taking stuff on Windows server that means that is the pointless you are able to start have have you you guys heard about such cases a few okay I mean brave souls aren't they brave souls yeah it means the I don't see future for the windows it's window yes in general it means that in the modern world uh I think that uh more there are probably no advantages of using Windows in comparable with Unix based systems uh and especially on servers uh even if you will compare what are the prices on cloud Solutions between the Linux systems and the windows based systems is completely without any sense to try with Windows but if you are willing to open your Microsoft Office on Linux then it keeps it's possible and even recently I found that Microsoft web browser called Edge has official Microsoft packages for Linux it works great you can use uh this J no GPT something Bing AI yeah Bing AI yeah yes because bigi Works only in Edge and Microsoft created Linux based Edge so probably Microsoft is going to the direction of Linux so they already make VLS so it means that maybe windows will be just on some playing stations for gamers even whole GPU intensive task right now like AI training is built on Linux it means that I don't know anybody who is training virtual a networks on Windows there are people who tries of course but but most of those most of those are people who are just Gamers and they have fun GPU and that's the reason why they tried it's not that they're building GPU workstation and this station is built on windows so thanks thank you very much was really nice giving this stuff
