Securing Grandine's Performance
Devcon·Tue, Oct 7, 2025, 12:00 AM
Our project focuses on improving Grandine’s performance and stability through targeted benchmarking and profiling. By conducting a comparative analysis with Lighthouse, we aim to identify architectural optimizations, especially those related to parallelization. Establishing baseline metrics is key to this approach, as it allows us to focus on refining critical areas within Grandine for optimal, efficient performance, thereby supporting the robustness of the Ethereum network.
Transcript
[Music] all right let's hear it for zarra BMA hi guys so we worked on consensus client performance profiling um both of us at the simar interest starting off with grandine and branching out a bit as you do during a work on a project um there's a couple of repos listed there two of them are on GitHub because that's where everybody else is and I myself prefer G laap so I threw that in um I'm going to hand it to Bulma who's GNA take the first bit okay hello hello okay sorry um my name is Messi and I'm here to make a pres presentation Although our projects shifted a little bit um the initial goal was to work on security and testing but due to some um uh I underestimated some things and so that was I had to focus on um performance analysis a comparative analysis between the cons client oh okay so okay sorry um so the initial goal was to do a security and reliability of grinding and then there was a shift in Focus which transitioned to to profiling multiple clients due to technical challenges which I faced oh okay sorry so um first of all I used so many performance security um tools the ones that stood out to me was a flame graph and then it implementing a timing metrix on grinding this is example of what a flame graph looks like uh although it's not quite readable because you have to like click in and then out to really view but okay so looking at this Stacks um grinding is actually performing very well from the stock overview although there's a drop down from the Tokyo run time which is at 61.79 per. which had a minus 35% drop down so the key observation here is there is a high Trend utilization on 97% in system Trends and initial setup significant drop down to 62 in Tokyo R time operations and then a consistent per performance across to tax management also multi trained worker execution showing similar patterns in different narratives is work okay sorry so um this is the 35% looking at the system trade level which is 97% then too one time which is which is 62% and then the which identifies a 35% drop down so the there is the the cliff there is kind of huge and then the high performance areas are trade initialization system level operations and then Co Trend management which gred pretty handled very well and then performed very well in that aspect so areas of Investigation okay a disclaimer this is a research project is still I will call it the research because so many things can change and then I might be wrong in some cases so areas of Investigation for me is a Tokyo TI sheding which is at 62% and then the sign corron time overhead and the tax profiling efficiency and um talking about the timing metrics which I had to implement for now I okay I was Sal suggested that I have to do a comparative analysis between um um grinding and then Lighthouse so for now I only have the data on grinding and I'm yet to produce the data for lious to properly do a comparative analysis but this is basically how the data looks like and I'm hoping to use this to do more investigation on um formal to do more analysis on formal verification and U foren creating a targeted foren and what's it called primtive analysis sorry and then hold on so the reason why I have to use use tools like the frame graph and the time um the timing metrics was to help me understand more about the consen client because I really underestimated what it is to force and build um a a forer for a conscious client because coming from a security a smart contous security background I thought it was the same perspective but I was proven wrong and the complexity was overwhelming but using this analysis and using this data I think I'll be able to Pro to pre produce and to come up with a strategy in in order to have a for to build a fora which has to do with um making it a targeted fora and also for formal verification because I understood that um there is no much um um um um work on formal verification on ethereum cons client thank you I can hand it over to you so small premise to the second part so as I mentioned we initially started working on grandine it was a slight shift in Focus transitioned to doing this profiling of multiple clients at least do what I focused on uh devop emphasis I decided to go there since I have no experience doing any sort of devops work and I figured having a couple of months seemed like a lot of time here that's a good thing to pick up so that was fun uh build my own server from secondhand part first first time doing that running proxmox um pretty interesting experience uh learned a lot challenges Hardware constraints I thought I had sufficient amount but turned out that maybe the minimal required specs listed are not entirely enough to run all the notes on a single uh server at least not the one that I decided to assemble um and regarding the outcome yeah well there's limited data uh the results are modest but valuable insight into running ethereum notes was acquired so let me tell you a bit about the setup that I used so I listed on the left side for you guys the consensus clients that are there and the versions of those that I used two of them don't have a version specified and that's because I didn't manage to run run them so they were not included I would like to include them at a later stage when I have the capacity to do so in terms of Hardware uh I have dedicated uh eight cores to each of these um Ram 32 GB and the dis space was two terabytes and VM Drive which is sufficient for all these clients to not be limited there in the setup I make sure that each of them running on a virtual machine is also is limited to this basically so there's no way that if one of the VMS isn't using all of its Ram that one of the others could uh use it because it's available this is normally something that Pro Mo supports but that of course doesn't allow for a proper comparative analysis um then the services that I also got to work with and got to learn about and nethermind of course I need an execution client and a whole bunch of other services that anybody who's done anything in devops probably knows so Prometheus obviously for metrics um I'm not going to mention them all but I will just say that these were a lot of tools that were new to me um quite interesting to see how many services you actually need to run such a such a node and a in a useful way so that you actually can monitor what's going on then I'm going to show a couple figures on some of the data it's a bit older data but uh so be it so this is the CPU usage that I measured over the course of About a Week a little longer uh for four of the consensus clients that I mentioned and now the question obviously is what does this mean and I don't have a crystal clear answer to you I do notice some things which you'll probably also notice straight away which is that grandine seems on the average to have the lowest CPU usage but there is somewhere a spike in the middle um I see that in a data figure in the next figure I'll show as well and honestly I'm curious what happened there but I simply don't know I probably need more data from the metrics alone it's not so clear it could be various things it might be something Network related could be something client specific simply don't have an answer for you there here I looked at the memory usage of the different clients um also you see the same spike in the same period but it's for a different client so yeah there's definitely something going on there and and I would say that my suspicion is that it has something to do with my server at that point since I see this weird pattern but like I said I have no straight answers for you and then the last figure I will show I think this is actually one of the things that is highly relevant I only show grandine and Lighthouse here because the initial Focus was more or less on grandine and well for comparison also taking into consideration Lighthouse so the number of pairs that you have is obviously very important uh and you can see that there are actually quite large spikes to the downside I have the limit set at 100 so that's why it looks kind of artificially kept there um but yeah if you lose a lot of pairs that's obviously not a good sign and well this seems to be fairly maybe not stable enough um or not as stable as you might want it so in terms of Outlook what I would like to do is refine and scale this node management so I now have almost automated setup for the provisioning of my VMS I would still need to switch to nyos or Taos to have it fairly immutable and item potent then I want to integrate e Docker this is something that I simply didn't have enough time to properly look into so so far what I've been using is sge uh well I started with simple scripts and running local binaries and eventually went on to use Docker compose but yeah I would like to move to kubernetes and I would like to also explore e Docker because I think it hoverers some of the things that I was lacking or missing in sge and then I only very recently that that is to say yesterday learned about the secret shared validator um so I'm going to look into this this is something in the direction that I was thinking I want to use this setup I want to use this setup myself with an agentic system because there I have a background um so I think it might actually partly at least overlap with this secret shared validator setup because I was thinking in that same direction and lastly I want to optimize the data queries because the idea of this agentic system would be to integrate those metrics so that the agent can proactively take the necessary measures for example manager node if you see that your peers are getting too low maybe you need to well do something to um to fix that um yeah that is all I have for you so then I only have a thank you slight left um obviously Mario and Josh it was amazing thanks so much for uh well all the time and effort that you guys spent and granting us this opportunity to be part of the EPF um yeah I think that's it so if there are any questions [Applause] then yeah uh hey thank you for an amazing talk uh I was wondering if the notes that you ran were validators or just notes no also validators also validators okay so it might be uh nice to see the difference between like a for each client what's different if they're a validator and if they're not yeah I agree there's lots of experiments you could do um really really lots more uh I would have liked to have done more of them but it's as BMA already said it's quite overwhelming because even when I thought I had all the pieces and sort of working and then you need to learn promql okay and then you learn a bit of promql and then you see all the metrics that are there for example Gra which I also used a little bit at somebody I just decided okay just cannot do it so yeah I agree there's lots more um and those those things would be interesting to look at you amazing thank you any other questions Mario um yeah um yeah I was wondering because you mentioned the setup with brm and then you want to automatize it setup with NEX it's really cool uh I'm just did you didn't really have in slides like some more specs about it like for example uh what are what virtualization are you using with Pro Mars like uh KVM or lxc or yeah KVM is what I'm using but I could explore maybe Alternatives I haven't explored any alternatives there so you think it's worth exploring uh just the Linux containers are good for servers as well uh and I was wondering like because proxo supports both and you just yeah you didn't have the details there so I was wondering like which one no this is true no this uh I use it's the lightweight version right so for Linux that's a that's a preferred option I would say cool thank you so much here all right thank you guys [Applause]
Automatic transcript — names and jargon may be misspelled.