New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Shadow Network Simulations by Daniel Knopik | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

In my EPF project, I implemented Ethshadow, a configuration generator for simulating Ethereum networks using Shadow, and used it to research improvements to the current state of PeerDAS and to estimate the effects of IDONTWANT on node bandwidth. In this presentation, I will present my findings and make a case for testing using Ethshadow. Speaker(s): Daniel Knopik Skill level: Intermediate Track: [CLS] EPF Day Keywords: Core Protocol, Layer 1, Testing Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

Hi everyone. Thank you. Um my name is Anja Kopik and I worked on Shadow Network simulations. Yeah. Um The usual clicker problems.

Thanks. Um yeah. So, in my opinion, network simulations are awesome because they allow testing changes without rolling out to public devnets or even testnets, which is slow and takes synchronization and time. And um small local devnets are possible, but with um a software called Shadow, we can actually run huge simulations with thousands of nodes. And also with actual clients like not any um like code written dedicated to the um simulations.

So, I kind of had a three-phase project. First, I wanted to prepare a tool that allowed me to easily set up these Shadow simulations with Ethereum clients. Um That is easier said than done because Shadow can be kind of finicky sometimes. So, it it took a while. Um yeah.

In the second phase, um I want to run experiments uh specifically on Pyrus and I don't want also known as Gossip sub 1.2. And in the third phase, I thought maybe if I saw in the experiments that this is actually a useful tool, um I wanted to polish it and actually did uh and such thus uh Eth Shadow was born. Um so to sum it up, there are like two main artifacts from my project, um, the experiment results and each other itself. In this presentation, I will kind of focus on on my results from Pier Dust and afterwards, uh, also show you slowly, um, or quickly, um, each other shadow.

Right. So, Pier Dust, um, I won't have time to explain Pier Dust fully, so I'm just going to assume there's, uh, some awareness of what that is. Um, few points, it's an update to scale blobs and the basic principle is that we want to split the blobs into parts, which we call columns, with every node taking custody of some of those columns. And there are some network challenge changes required for this to make sure that the nodes, um, properly, um, distribute, uh, the columns across the network and also so-called super nodes can reconstruct the blobs if they have enough columns. And right now, there's like 128 additional subnets, um, in gossip sub, uh, which all need to be, uh, covered by, when a node wants to be proposed, right?

We need to get every column out, so, yeah. Okay, so I decided to run some simulations to help the research and development effort. Um, keep in mind that all re- results that I present here are like a few months old at this point, so the situation might be, uh, better at this point. So, my simulation setup involved 1,000 nodes with Reth, Lighthouse and Lighthouse validator client on each, uh, with 4,000 validators and in each experiment, I simulated 45 minutes of simulated time. Um, for a large simulation like this, we want to run on some server that is sufficiently large so the simulations fast enough for us.

And I kind of experimented around a lot to make sure to get like get nice performance going and kind of settled on a certain instance type. Um, you can see it there if you want to run your own simulations with this size. And my simulations took around of 4 hours per like 45 minutes of simulated time. And so the the cost of per simulation came around at $10. dollars.

Right. Um, but if you want like also run smaller simulations, you can you can also do this on your laptop. Like you don't need such a beefy server. This is just like my setup for the beta simulations. Um, so first I just tried running it on the current implementation in Lighthouse.

And the simulated networks quickly fell apart because they lost sync and cannot really sync back up again. Um, and that was kind of similar what was seen on devnets at the time across basically all clients. Um, so I thought about how do I measure performance and I had several metrics but right here I will focus on what I call score, uh, which is how many slots can um, over 66% of the network uh, stay in sync. And with the unmodified client with the default configuration, zero. Like immediately after we posted blobs to the network, um, we lost sync.

Um, but I also noticed that if I designate every node as a super node, so custodying all the columns, not just eight, uh, or four, um, then, um, the network stayed stable until the end. So, I had like the suspicion that something is wrong with um the distribution of uh columns when we just, you know, don't broadcast all the columns at all time to everyone. And also that might be something with supernode reconstruction of uh the blobs might be wrong. Yeah. Um So, and yeah, the investigation showed that the problem is that if a node could not send all the data out, it would not retry that at all, causing some columns to get lost, and custody nodes uh for those columns would just not deem the data as available, and would just not accept the block and lose sync.

And they could not catch up, I guess. The supernodes were not not enough for that. So, how do we fix this? I did some simulations where I varied the specs, and um I did a lot of different um variations. I will focus on three here.

Uh we could increase the number of peers because if we have more peers, the likelihood that we actually cover all the columns uh when proposing a block um with our peers uh is higher. So, yeah, uh the base back or the base configuration in Lighthouse uh looks for 100 peers. 150 peers uh scored uh 10 with the metric I uh just mentioned. 200 uh already uh 56, and 300 survived the whole simulation. So, we could see that, confirming our theory, this improves the situation.

Next up, I increased the number of columns covered per each peer, which should have the same effect, right? Because the total number of covered columns also increases with uh changing this. Um increasing from four to eight scored 10, and 16 already survived the whole simulation. Finally, uh there should be an emoji here. The the the you know the the diagonal face.

Um All right. Increasing the number of supernodes, which should reconstruct the um full blobs like all comes if they have half or more. Um increasing to 10 scored zero to 25 of thousand, uh right? I scored two and 75 also two. And I don't But I didn't test any further because I want to kind of keep it in a somewhat realistic scape, and I'm not sure how many supernodes there will be, but we shouldn't, in my opinion, set the assumption too high.

All right. Um And if you want to know more about that, um there are my weekly reports where I go into a bit more detail. So, let's uh move on to the state of like the tool I developed. Um because seeing these simulations, um I thought this is hey, this is kind of useful. I have like a an iterative approach and can evaluate different configurations, so I, yeah, spent some weeks in the end to make it available for everyone.

And the source is available on GitHub at uh ethereum/ethshadow. Uh right now, there's like first-class support for Geth and Lighthouse, and there is some experimental uh support for Reth and Teku. Um there is some documentation left to be written. Um for basic use cases, there's like good uh documentation already, so you can check it out. Um and I think in some cases, usability can be improved.

Um the hard part is that Shadow needs a lot of work for all clients to work in Shadow. The reason for that is about the technical Basically, we need support for more Linux system calls in Shadow. Um and this is a lot of effort. Unfortunately, I um kind of won't have much uh time. So, I hope that in the coming uh weeks and months I can at least help on the side a bit to maintain this, but I hope to also get the attention from from the core devs to help me uh or help us all um hopefully soon have each shadow for all the clients.

Okay. Uh there are some things to be said. Uh first of all, thanks to my mentors, Eireann Manning from Sigma Prime and Bob from Ethereum Foundation Research. And also thanks for to João Oliveira and Jimmy Chen from Sigma Prime who supported me with some uh of their time. And thanks to Anton Nashatyrev from ConsenSys.

Uh he uh spearheaded uh the technical support in a couple uh last couple of weeks. Uh thanks to EPF organization uh to Josh and uh Mario. Thanks to the Ethereum Foundation for hosting the EPF. And uh thank you for to all the fellows. It was really pleasant uh and great to work with you all.

And finally, I'm going to plug my talk tomorrow or our talk tomorrow um simulating an Ethereum network at scale, which will go more into detail on how you can run these simulation yourself. So, maybe see some of you there on stage one at 10:01 p.m. Thank you for your attention. All right.

Any questions for Daniel? Um hi. I'm just uh a couple of questions. One is um when you're making all of these changes to do these simulations, are you actually like updating the client code or is it a kind of like network configurations? Um yeah, this is like the great advantage to change these or to test these changes, I actually can change just the client.

I don't need to develop anything just for that. I change the actual client, but for like the things I showed you, I think I only needed to vary the configuration a bit, right? But, I also did some changes that actually tried to fix these underlying problems uh in light at the cell. Cool. And then, I guess my other question So, are you starting simulations from like a genesis state, or is it like a fork of mainnet, and can you do forks of like existing networks?

Well, um I I start from a genesis state. My tool supports this. It It just does this for us, uh which is nice. And forking mainnet I mean, I'm aware that there are shadow forks. They are unrelated to the name of uh the ship simulation tool.

Um I haven't tried it. Uh the problem is that the genesis state is quite large, and this might actually be really hard to simulate like on a single uh node, right? We We have all these simulations running on a single server, and that might be hard to have multiple clients in parallel uh working on that. So, I guess it's not really feasible. So, I was wondering like once the simulation completes, can we see like consolidated metrics on the node or something like logs, something like that?

Very good question. Uh this is how I evaluated uh those simulations. Um I kind of forgot to mention that here. Um the simulation um it's very really really easy to add a Prometheus node in the simulation configuration, and that automatically gets configured to pull uh the metrics from all the lighthouse clients. And in the end, we have like one huge Prometheus database, which allows us to just pull up some Grafana uh dashboards or any other analysis.

Also, we have all the logs, so we can also look into those if there are any like specific nodes that seem to be acting strange or something. Yeah. Right. Um you can also like run a node on a data directory that has been generated by the simulation afterwards. It's a bit weird because uh the simulation time always takes place in the year 2000, so it's really old data to the node if you start it uh right now.

But you can run that node and you can attach like a block explorer or the explorer for example by the PandoOps team to it. It's just a bit wonky because of this timing issue. Yeah, you basically the simulation shadow starts the simulation time always at the 1st of January in the year 2000, so every time you do a Prometheus query or um yeah, look into the block explorer, you have to keep in mind that you have to act as if it were the year 2000. Yeah. Any final questions?

There are some up there. Oh, we got one more. Uh how is it different from Kurtosis? Okay, in a nutshell uh Kurtosis um you cannot run networks at of this scale on a single server with Kurtosis because Kurtosis is not capable of like pretending to the processes that time is running faster or slower than it actually is. Basically like um having a a a separated simulation time from real time.

Um yeah, I would say that this is the main difference that you can scale way higher with uh Shadow. All right, thanks Daniel.

Automatic transcript — names and jargon may be misspelled.