A toolbox for monitoring the health of the Ethereum P2P Network | Devcon SEA
Devcon·Tue, Oct 7, 2025, 12:00 AM
Monitoring the P2P layer of Web 3.0 networks is extremely critical for the healthy operation of the system, but has been overlooked for quite a while. ProbeLab has developed a suite of open source tools to monitor closely the journey of blocks and messages in the Ethereum network, as well as the operational details of Ethereum nodes. The target is to assess the health of the network as a whole.In this workshop we’ll walk through and demo the details to enable others to benefit from our tooling. Speaker(s): Yiannis Psaras, Dennis Trautwein Skill level: Expert Track: Core Protocol Keywords: Protocol Design, Tooling, Network State, gossipsub Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/
Transcript
[Music] hi everyone welcome uh I am Janis I represent the pro blob team and uh today I'm going to talk about a set of tools that we have built to monitor the health of the ethereum peer-to-peer network uh why we're doing this um because we believe that things generally might be working but not necessarily in ways um in the way that we think they do so we think that we should generally be looking deeper into um how things are working and we argue that the peer-to-peer layer of the um of the web 3 ecosystem of several blockchains has received a lot less attention than it should have there is a ton on chain metrics and dashboards but very little has been done on the peer-to-peer layer which is under everything uh under everything else so we argue that if something breaks at uh that layer then everything above it all the great stuff that people are building and presenting is just going to come crushing down so um what if I had to say in one sentence what we're doing is that we're carrying out rigorous measurements on the peer-to-peer protocol stock in order to make the network layer better faster and safer uh and we think that's what is needed in order to guarantee a robust layer below um all of the other stuff that is building on top now over the years we have built a set of tools uh starting from more uh kind of high level topology and network structure um items and going deeper into critical components monitoring and um uptime monitoring and so on and then going deeper into more detailed and sophisticated things like uh data availability challenges that the community is uh trying to overcome at this point uh I'm obviously not going to cover everything uh but just going to focus on a few of those um tools to go through and let you you know give you an idea of what they are and the related uh points so starting with network topology you might have heard of or used the nebula crawler which is a very welln crawler by at this uh by this point uh it started off as uh a crawler for the ipfs network but has now uh expanded and supports many other lip2p based networks but also dis V4 and dis V5 networks such as uh the ethereum consensus and execution layer uh the uh one thing that may nebula um you know come apart is that it's not only crawling the network it's also monitoring the liveness of node so it's going and you know um crawling networks finding peers and then pinging them every so often after a while um like continuously and therefore we can find things like uh P CH and so on so uh let's start with a quiz to see what we can find out um from uh from things like a simple crawler um and uh all of the other things that are built on top so which one uh which ethereum consensus layer client do you think is um least dependent on centralized Cloud infrastructure yeah others okay that was correct actually the last one Nimbus is the one um well done to to the team so uh if if we if we see so at the top here it's not shown very well but it's Lighthouse prism um and load star and at the bottom is Teo grandine and Nimbus so if you see it in terms of percentages obviously Nimbus is uh the least dependent but of course it's not the most popular uh client so it depends a little bit on how you uh how you do it the the the blue line here is the data center those that are deployed in data centers and the orange one is those that are um have got custom deployments um so you can find other things like what's the most popular chain on uh optimism any ideas anyone sorry Bas no this is a pretty recent result and we found that is actually uni chain um by far uh so yeah this is pretty recent if you if you think um we're missing something or or not we can uh come talk to us we can we can argue and uh have another look later uh to figure out if we're correct or not um so yeah this is uh these these are things that you can uh figure out if you look a little bit deeper into the network uh and that's what we uh we're actually doing so carrying on again on the network topology we have recently uh built a tool called ANS um which uh you know like ants when you release them and when they they're out they're covering um a big footprint so what we're doing is that we want to release ants in um DHD structures so that uh they can monitor specific things specifically DHD client requests um so we we spread ANS um on every approximately every 20 DHT uh server nodes and that way we're able to gather information from all the requests from the network the two is automat automatically adjusting so if if the network size doubles then the number of ANS that are going to be out in the network is also going to double and we're currently uh supporting the Celestia network but soon uh there there's going to be more so briefly how it works is that you see uh we we can assume this is the DHD keyspace uh these are reg regular DHD server nodes uh the green ones are uh the ons that we are putting in the network uh as I said roughly every 20 DHD server nodes and therefore then what happens is that when a DHD client comes in uh in order to do several processes that he needs in the network among others uh update its rooting table it's going to ask other nodes around so we want to make sure that one of these green nodes is among the 20 um peers in the network where the DHT request is going to end up to so in that way we're kind of uh it's hanot of sorts that is gathering requests um now with this uh as I said it was uh great to see how you know we were able to figure out what's the light node population in the Celestia Network um we don't have results yet but it's soon going to be on Pro blob.io so if you want to um know more about that that's that's the the thing to monitor or you can come talk to us afterwards um now going into protocol performance um we have recently built a tool called Herms which is basically a gossip sub listen listener and Tracer uh the measurement goals of building that was to figure out metrics related to message propagation latency number of dou um messages that a node is receiving um the validator node bandwidth consumption and several others so in doing that we didn't want to run a full ethereum node for several reasons what we instead wanted to do is have a lightweight lip2 be host which would spin up establish connections stay connected and then participate in the network in this way um it's tracing Herms is tracing all of the events so from the lip2p host events like connections and disconnections to gossip sub Trace events which were of particular interest to us uh like grafts and prunes and like in the off nodes in the local mesh Network subscriptions I have and I want messages for the gossiping function of the protocol and then go deeper into peer scoring and and things like that so we were able to to gather all of these data and uh this helped us look into lots of um important metrics for the network for uh a protocol for gossip sub that is very critical right it's uh as you might know it's transferring blocks and messages um in the ethereum blockchain and other blockchains as well uh this is roughly what what the architecture looks like but I'm not not going to stay in in this slide slide because uh that's just one way of doing it there might be several others and even this might evolve over time but this is what we used uh for the um for deploying Herms on um on the ethereum network now um briefly some results so um we wanted to see the number of duplicate messages and in this plot there you see the number of duplicate messages on the x-axis so it goes from 0 to 10 and the message ID the CDs uh of messages on the Y AIS so if you um if you look at the green line it's representing the blocks topic um sorry the blob topic and we see for example that around 55% uh of messages are being uh delivered to the same node four times or less so on the one hand you know we like seeing this result one might say four times is is quite a lot on the other hand you know we we should not expect to have you know zero duplicates because this is not a client server architecture here it's a p-to-p network and having duplicates in many cases um helps in in in case for example of an attack where someone is trying to delay delivering messages or eclipsing a node uh so it's good to have some duplication but we're seeing that as you know uh when we're going to six or eight duplicates per message that might be a little bit too much so um we we figured out that things are working well but there is still some space uh for optimization um on another bandwidth related um metric that we've seen this is um on the on the x-axis of the uh message types um and on the y axis is the percentage in terms of um kilobytes from the total and we we found out that I have messages represent 33% uh of the total bandwidth of an mod uh which is quite a lot and there might be again some of it is expected but um there is clearly some space for optimization um you'll find all of the reports that we've done on E research uh if you go it's uh all of these studies you can scan them but also under the networking category um they're pretty recent the last few few months so you can uh take the full picture of all the things that we've uh managed to um to find um now moving on to another um to another tool that we have uh recently built it's called ukla Uh a network bandwidth monitor uh we might have to change the name but uh that's what we have for now uh so the methodology is that um we we as we run nebula and we have a list of online peers we then deploy Herms and try to connect to those peers and keep uh stable connect connections and then eventually we're trying to download a carefully chosen volume of data from each node so that on the one hand we do not overload um the node that has got other things to do like these are main nodes that are uh doing their thing and at the same time we we are able to kind of momentarily saturate the bandwidth of the node so that we know how much bandwidth uh they've got available um we're seeing in the in this graph here on the uh on the x- axis is the the slot duration from 0 seconds to 12 and on the Y AIS is the bandwidth in megabits per second so interestingly we're seeing that there are some uh like dips in the um bandwidth availability the Blue Line especially for um the blue line which represents Cloud nodes um whereas the the orange one is for not non-cloud nodes so seeing these dips there means that there is a little bit of extra bandwidth availability uh which could be very interesting for engineering teams if they want to for example push some more data out during some particular you know time within the slot then this could be chosen um and the this I believe it's going to become very helpful as um you know as we're going towards piras and the blob count increase and all of those challenges where you know it's not exactly clear where bandwidth availability can can come from in order to to accommodate uh big a blobs so on the one hand Nitty Nitty nitty-gritty details on the other hand um we believe that measurements is not really an end but only a means to an end which is you know to go towards protocol optimization and network architecture optimization and that's how um um we're helping here so um yeah to recap um quite a few tools we've uh covered uh some of them on and again we're going from um Network topology and structure into more detailed things and in most cases we end up um having to use several of our tools as we as we go along so as I mentioned having nebula gather a list of peers then using that to inform other tools to go and connect with those peers and then go from there deeper and deeper uh and we believe there is there's lots more that can be um can be done so uh looking ahead we obviously are collecting data for a number of networks for a number of years now so we have a kind of data warehouse of sorts which we want um to make available to the community so we're building an API um we know that other um other teams are interested to get the data and deploy into their own models to do energy sustainability metrics and risk modeling and and things like that so this is under active devel development right now uh we want to integrate more networks um as you know with some of them on the pipeline already uh we want to repeat an nut traversal study that we have done a few years ago on Li P2P to uh figure out the capability of the nut hole punching component and then um more generally focus on the challenges to deal with date availability sampling on ethereum and other networks so if you head to Pro blob.io you're going to find out weekly reports for uh the the the networks that we're currently monitoring uh you're going to find other reports like like blog propagation latency for the ethereum network and then uh links from there to our methodology and all of the tooling um yeah and as I mentioned of course some of the results St studies uh results of the studies are on E research um yeah and prob. network has got some more information about uh our team um yeah and now thank you very much I'll now pass it on today is to do uh a little demo of one of the tools the the bandwidth measurement monitoring tool to see what bandwidth looks like uh on this uh Wednesday afternoon all right hi everyone my name is Dennis I'm also from the pro blab team okay [Applause] yeah all right can you see my screen yeah there we go okay it's loading sorry ah here we go cool so Jan's introduced you to ukla already and to spice it up a little bit I wanted to try uh a demo to show what the data looks like that we gather um I mean this data won't be representative of cors because we're running it from this uh Convention Center Wi-Fi um but still I think it should be interesting to you uh to see what data we We Gather so if we run Uka uh this is our band withd measurement tool as you have heard uh you you can pass in if you um line parameters for example in this case we querry 10 peers from our nebula database which are peers that are online in the uh ethereum Network we request from each of these 10 Piers two blocks and if it fails we do three retries and because we are running quite a few of these um ukla runs every day um you can also give it a tag from where you have run it for example we are currently running it from us East one but you could Als in AWS but you can also run it from different uh ge locations for example and I just give it a tag of flip of of Defcon day Defcon demo or let's leave it thep for now um so right now oh okay of course the demo demons it's always the same sorry so we are running um a full prism note in the uh on one of our inst instances in the cloud and Herms as youve known is part of of ukla which is running underneath um so which we build upon it needs to connect to one of these full nodes uh so I'm just forwarding the port here let's try it again so it's launching the Herms note and I want to emphasize Herms is um quite modular so if you require for your network for example bandwidth measurements Herms could be extended uh to also work for your for your network if you have bandwidth measurement requirements so let's see okay so basically what what it does I can talk talk over it um it will yeah it will basically just connect to the notes download the blocks and maybe in the meantime I've have only two minutes left I have run it earlier so we might take be able to take a look at okay it's all not working of course so meantime just yeah I encourage the audience to connect to Mir cut app and you can ask questions hang on let's see okay I hoped it would work on and I I practiced a lot of course it's not working let's let's try it again okay may maybe you can ask questions and I try to run it in the meantime yeah you will answer the questions yeah so um ah here we go it's working so it's it's it's connecting to notes downloading yes thank you very much so it's now connecting to notes downloading blocks and um measuring the bandwidth that these nodes have available to get um a distribution of what bandwidth is available in the ethereum network which then would inform protocol decisions around um for example data availability s or block sizes that we could um increase to some certain degree uh and so on um The Next Step would have been that we look into the database but it looks like my computer is frozen here so maybe you should maybe that's already enough to Showcase that it's working for now and maybe go on with some questions thank you okay yeah we have a first question and how many noes do you need to run yourself to get a complete picture of the network uh okay I I don't know for uh a complete picture of the network I guess is that for the ANS if uh complete picture of the for the for which tool basically I don't that's why I don't yeah if the person is in the audience okay so can you pass the micro thank you yeah I was just wondering more in general terms for all the tools everything thing to get all the data you need like besides the crawler for ANS or for anything like that how many noes you run yourself in order to get access to the networking and data or can you just use public nodes and connect to that right so um it depends on the tools like uh we we do have uh for example the nebula CER is running 24/7 for from um uh for several networks we have some other tools like um to monitor the DHD up time of the um ipfs DHD Amino DHD and that is running from several locations from Seven locations around the globe I think again running 24/7 um yeah I don't know if the answers uh the question for for the ANS architecture in particular for the Celestia Network we're running something between five and seven nodes I think it doesn't it doesn't need more because it can cover more space yeah that makes sense thank you very much yeah uh let's let's go through a couple more questions and then return to the demo yeah Co so um can I run these tools within kosis local networks I guess the answer is yes here depend so I'm not familiar with kosis and the networking set that that it's using um yeah uh hey uh yeah you can configure ketosis so that it exposes the ports outside of just the network that in Docker uh so it essentially exposes it on the host yeah uh and then yeah all of the stack will work yeah I think and another interesting question is if these tools can be used in a customly P2P Network the second question yeah generally the answer is yes uh it's not um as like most of the tools that we're building are for Li P2P based Network Works uh so it's going to be a lot easier than building a new tool from scratch uh that being said uh something things depend on like the the network characteristics and how the whole um you know protocol is uh is set up uh it's not a matter of just changing one parameter and adjusting to another network but we have done I think a very careful job of making it easy to integrate more networks okay and one more question what stops someone from using Uka to grieve the network uh which one ah first yeah um the right now it's not open source I think that's the easiest answer that's how we stop yeah uh okay I think we don't have much time left so let's just transition back to the demo yeah I just want to briefly show the data that we it's a good question though this one because we are planning to make it open and source so we'll think about it yeah Okay cool so so you can see the database uh so as you've just seen uh it has run and you can see that um there are two studies now because I run it be right before the talk and it was working and uh so the second one here is the second study that we just started here you can see the tag that we gave uh to it and then each connection to the to an ethereum node um results in a visit in a bandwidth visits how how we call it here and um yeah here you can see all the data that we together here so how um in which time of the slot did we request the data so you can imagine that in different times of the slot the bandwidth might differ because you have to do other duties here um if the connection was successful which blocks we requested from that other node um the um the compressed number of bytes but also the uncompressed number of bytes the btes that that we sent over the wire so with all the uh enveloping and so on and how long it took and then the resulting megabytes uh per second and yeah the round round time and how many retries we needed and so this goes even Beyond just the plain bandwidth me measurements um yeah so there's even more data that we can use to inform protocol decisions for example which is one of our main goals so thank you okay okay thank you and yanis if you want to say some uh last remarks U uh no yeah thank you thank you very much thank you danis and yanis for amazing tool to analyze that Network I think now it's time for
Automatic transcript — names and jargon may be misspelled.