Shadow Network Simulations
Devcon·Tue, Oct 7, 2025, 12:00 AM
In my EPF project, I implemented Ethshadow, a configuration generator for simulating Ethereum networks using Shadow, and used it to research improvements to the current state of PeerDAS and to estimate the effects of IDONTWANT on node bandwidth. In this presentation, I will present my findings and make a case for testing using Ethshadow.
Transcript
[Music] hi everyone thank you um my name is D kopic and I worked on Shadow nworting relations yeah um the usual clickup problems thanks um yeah so in my opinion Network donations are awesome because they allow testing changes without rolling out to public death Nets or even test Nets which is slow and takes synchronization and time and um small local definites are possible but with um a software called shadow we can actually run huge simulations with thousands of notes and also with actual clients like not any um like code written dedicated to the um simulations so I kind of had a three-phase project first I wanted to prepare a tool that allowed me to easily set up these Shadow simulations with etherum clients um that is easier said than done because Shadow can be kind of finicky sometimes so it it took a while um yeah in the second phase um I want to run experiments uh specifically on pias and I don't want also known as gossip sub 1.2 and in the third phase I thought maybe if I saw in the experiments that this is actually a useful tool um I wanted to polish it and actually did uh and S thus uh e Shadow was born um so to sum it up are like two main artifacts from my project um the experiment results and E Shadow itself in this presentation I will kind of focus on on my results from pias and afterwards uh also show you slowly um or quickly um e Shadow right so pias um I won't have time to explain pias fully so I'm just going to assume there's uh some awareness of what that is um few points it's a update to scale blobs and the basic principle is that we want to split the blobs into Parts which we call columns with every note taking custody of some of those columns and there are some Network CH changes required for this to make sure that the nodes um properly um distribute uh The Columns across the network and also so-called supernes can reconstruct the blobs if they have enough columns and right now there's like 128 additional subnets um in gosip sub uh which all need to be uh covered by when a note wants to be proposed right we need to get every column out so yeah okay so I decided to run some simulations to help the research and development effort um keep in mind that all results that I present here are like a few months old at this point so the situation might be uh better at this point so my simulation setup involved 1,000 notes with r Lighthouse and Lighthouse Val on each uh with 4,000 validators and in each experiment I simulated 45 minutes of simulated time um for a large simulation like this we want to run on some server that is sufficiently large though the simulation is fast enough for us and I kind of experimented around a lot to make sure to get like get nice performance going and kind of settled on a certain instance type um you can see it there if you want to run your own simulations with this size and my simulations took a round of 4 hours per like 45 minutes of simulated time and so the the cost of per simulation came around at uh $10 right um but if you want like also run smaller simulations we you can also do this on your laptop like you don't need such a PV server this is just like my setup for the P simulations um so first I just tried running it on the current implementation in Lighthouse and the simulated networks quickly fell apart because they lost sync and cannot really sync back up again um and that was kind of similar what was seen on Def Nets at the time uh across basically all clients um so I thought about how do I measure performance and I had several metrics but right here I will focus on what I call score uh which is how many slots can um over 66% of the network uh stay in sync and with the unmodified uh client with the default configuration zero like immediately after we posted blobs to the Network um we lost sync um but I also noticed that if I designate every node as super no so coding all the columns not just uh eight uh or four um then um the network stayed stable until the end so I had like the suspicion that something is wrong with um the distribution of uh columns when we just you know don't broadcast all the columns at all time to everyone and also that might be something with super node reconstruction of the blobs might be wrong yeah um so and yeah the investigation showed that the problem is that if a node could not send all the data out it would not retry that at all causing some columns to get lost and custody nodes uh for those columns would just not deem the datas available and would just not accept the block and lose sync and they could not catch up I guess the supernes were not uh not enough for that so how do we fix this I did some simulations where I varied the specs and um I did a lot of different um variations I will focus on three here uh we could increase the number of peers because if we have more Pi the likelihood that we actually cover all the columns uh when proposing a block um with our peers uh is higher so yeah the base back or the base configuration in Lighthouse uh looks for 100 Pi 150 Pierce scored 10 with the metric I just mentioned 200 uh already uh 56 and 300 Sur the whole simulation so we could see that confirming our Theory this improves the situation next up I increased the number of columns covered per each Pier which should have the same effect right because the total number of cover columns also increases with uh changing this um increasing from 4 to 8 score 10 and 16 already so we have the whole simulation finally uh there should be Anem Emoji here the the you know the the diagonal face um right uh increasing the number of supernotes which should reconstruct the um full blobs like all columns if they have half or more um increasing to 10 scored zero to 25 of thousand uh right scored two and 75 also too and at that point I didn't test any further because I want to kind of keep it in a somewhat realistic scape and I'm not sure how many supernes there will be but we shouldn't in my opinion set the Assumption too high all right um if you want to know more about that um there are my weekly reports where I got into a bit more detail so let's uh move on to the state of each Shadow like the tool I developed um because seeing these simulations um I thought this is hey it's kind of useful I have like an iterative approach and can evaluate different configurations so I yeah spent some weeks in the end to make it available for everyone and the source is available on GitHub at ethereum Shadow uh right now there's like first class support for GU and lighthous house and there is some experimental uh support for re and tiu um there is some documentation left to be written um for basic use cases there's like good uh documentation already so you can check it out um and I think in some cases usability can be improved um but the hard part is that shadow needs a lot of work for all clients to work in Shadow the reason for that is are technical basically we need support for more Linux system calls in shadow um and this is a lot of effort unfortunately I um kind of won't have much uh time so I hope that in the coming uh weeks and month I can at least help on the side a bit to maintain this but I hope to also get the attention from from the C devs to help me uh or help us all um hopefully soon have each for all the clients okay uh there are some thanks to be said uh first of all thanks to my mentors Aran Manning from Sigma Prime and pup from etherum Foundation research and also thanks for to J oliviera and Jimmy Chen from Sigma Prime who supported me with some of their time and thanks to Anon nurv from consensus uh he uh spare headed the tech support in a couple last couple of weeks uh thanks to EPF organization uh to Josh and uh Mario thanks to etherum Foundation for hosting the EPF and uh thank you for to all the fellows it was really Pleasant uh and great to work with you all and finally I'm going to pluck my talk tomorrow or our talk tomorrow um simulating an ethereum Network scale which will go more into detail on how you can run the simulation yourself so maybe see some of you there on stage one at 10 1 p.m. thank you for your attention all right any questions for Daniel um hi just a couple questions one is um when you're making all of these changes to do these simulations are you actually like updating the client code or is it kind of like Network configurations um yeah this is like the great advantage to change these or to test these changes I actually can change just the client I don't need to develop anything just for that I changed the actual client but for like the things I showed you I think I only needed to varry the configuration a bit right but I also did some changes that actually tried to fix these underlying problems uh in L itself cool yeah and then I guess my other question um so are you starting simulations from like a Genesis state or is it like a fork of main net and can you do Forks of like existing networks well um I I start from a j St my tool supports this it it just does this for us uh which is nice and forking main net I mean I'm aware that they are Shadow Forks they are unrelated to the name of uh the simulation tool um I haven't tried it uh the problem is that the Genesis state is quite large and this might actually be really hard to simulate like on a single uh node right we have all the simulation running on a single server and that might be hard to have multiple clients in parallel uh working on that so I guess it's not really feasible so I was wondering like once the simulation completes can we see like consolidated metrics on the Node or something logs something like that very good question uh this is how I evaluated uh those simulations um I kind of forgot to mention that here um this simulation um you it is really really easy to add a Prometheus note in the simulation configuration and that automatically gets configured to pull uh the metrics from all the lighthouse clients and in the end we have like one huge Prometheus database which allows us to just pull up some grafana uh dashboards or any other analysis also we have all the locks so we can also look into those if there are any like specific noes that seem to be acting strange or something yeah right um You can also like run a note on a data directory that has been generated by the simulation afterwards it's a bit weird because uh the simulation time always takes place in the year 2000 so it's really old data to the note if you started right now but you can run that Noe and you can attach like a block Explorer or the Explorer for example by the e p Ops Team to it it's just a bit wonky because of the timing issue yeah basically the simulation Shadow starts the simulation time always at the 1 of January in the year 2000 so every time you do a Prometheus query or um yeah look into the block explore you have to keep in mind that you have to act as if it were the year 2000 yeah any final questions there are some up there oh we got one more uh how is it different from curtosis okay in a nutshell uh kosis um you you cannot run networks at of this scale on a single server with kosis because kosis is not capable of like pretending to the processes that time is running faster or slower than it actually is basically like um having a separated simulation time from real time um yeah I would say that is the main difference that you can scale way higher with uh shadow all right thanks Daniel [Applause]
Automatic transcript — names and jargon may be misspelled.