# EPF - Nethermind/IL-EVM by Siddharth Vaderaa | Devcon SEA

- Speakers: [Siddharth Vaderaa](https://streameth.org/speakers/siddharth-vaderaa)
- Channel: [Devcon](https://streameth.org/devcon)
- Date: 2025-10-07
- Duration: 11:40
- Watch: https://streameth.org/watch/yt-9UhpqUzsEJE
- YouTube: https://www.youtube.com/watch?v=9UhpqUzsEJE

## Description

This talk will discuss my EPF work on Nethermind's IL-EVM project, which included developing tools to analyze EVM execution patterns, writing  (optimised) opcode and top pattern implementations, and conducting and writing tests.

Speaker(s): Siddharth Vaderaa
Skill level: Intermediate
Track: [CLS] EPF Day
Keywords: Core, Protocol

Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon
Learn more about devcon: https://www.devcon.org/
Learn more about ethereum: https://ethereum.org/ 

Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more.

Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. 
Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024.
Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

## Transcript

[Music] right so this is my project uh EPF uh NE mind I evm IL evm is a evm optimization project from nethermind uh in uh seha this is sort of the lose timeline of uh of the project so week six and8 were sort of done on some warm-up tasks to kind of get a feel of the base uh week 9 and 12 was done uh was spent on doing some research for the upcoming task which is basically Gathering engr stats the core focus of the project was doing the stat analyzer implementation which happened between week 13 and 16 and then week 17 and 18 I was basically running the stat analyzer getting the end GRS and then actually doing the M implementations uh for patent detection mode for the and last I was just basically doing bug fixing and also just generally looking at ilm and bugs and I found like five off codes that had problems and I fixed those um so right so the first part was basically it was um it was my first task um I just joined I I just started the project and uh I think it was the first meeting with the mentor and I basically was uh the mentor asked me do you know where I was because the mentor was of sort of focusing on eof implementation so nobody was working on ilm I just sort of took the task that was sort of remaining and I started working on it and then my mentor joined in and that was sort of finished the next task that I did was some cod DB stats again this is a warm-up task where I just Tred to get some engram stats from the database so these are not execution stats uh the implementation was very slow it was done over the weekend again it was a warm task so now we kind of come to the you know the oh this is the research and algorithm so there was a lot of research and literature review that was done for basically Big Data analysis because we're dealing with the execution stats for MGR which is actually like a lot of data and these were some of the papers that were run uh that that I that I read uh of Nota basically the heavy Keepers and also sliding sketches uh in time zones for data stream processing because these are both single pass algorithms because we wanted to implement an algorithm that is single pass for for this amount of data and and is efficient so uh after a while we just sort of U with some discussions with the mentors we just settled on simple cment sketch uh the basic principle of the count sketch is that you have um sort of dhash functions and then you have like let's say w buckets and whenever an item comes that gets hashed and it's put into one of these buckets so for example here you can see just data for one item so in this case the true Count will be the lowest count that is one 25 would be an overestimation that means there's a collision so but because we have many hash functions if you take just the minimum of the count you can actually get the true Count uh adjusted with errors Etc uh for the data that we need so uh additionally uh you have you have these bounds of Epsilon and Delta uh that that control the probability and the error uh that the data structure has and you can actually configure the stats analyzer to use both width uh depth uh Epsilon and uh and Delta so now for the main thing which is building the stat analyzer uh okay the first part was encoding the end GRS so to C the uh the task was given is that basically the uh we just need to find uh 2 to7 um op code patterns like 2 to 7 the size of the engram is 2 to 7 and the way to basically do it and this could work for like up to n engrs actually uh it's only limited by the data type but basically as the engrams are coming as the op codes are coming in we just basically left shift a op code and then we or it with the next op code and we encode it into a long value right so because a long value has eight pites so it can actually encode eight off codes and if you have like used uh one of the data types that is common in clients is like you have a 32 byte data type like un 256 uh so technically it could go up to like uh I don't know 32 uh 32 op 32 n g like size of the pattern it can actually technically go up uh wait right then comes tracking the engrams so instead of like tracking each engr separately like of 2 3 4 5 6 7 what we can do is we just uh get the nram the long value that we have we we find the NR that is the maximum uh of the size like for example if we at the size three right and you can actually have 0x f f FF which is size two that that's the maximum because you cannot have anything over this will be a size three and then we can actually uh select that out and basically if you have like uh an engram like pop pop ad data you can actually just iterate over those that long value and actually get these three NRS that don't have to be tracked individually and saved anywhere uh and they just basically you have these sort of ephemeral uh long values that just simply go into the CM sketch for accumulation yeah iterates on the back the Bas the core of the stat analyzes basically it's aray of these um count pin sketches and you have a top k q you iterate over the bite code you encode the engr you culate the counts in the CM sketch uh also a new sketch is provisioned um based on uh what our error threshold is like what is the max error we able to tolerate so we you can configure that and then new sketches a provision based on that uh it provides the top K patterns and uh provides the error and confidence for the stats so you can actually have very performances for this so you can actually make it really fast and less accurate or slower and more accurate you have that capability uh right so then how do we gather the data the the data is done as a trace it's a trace plug-in uh it's straight from the evm you can configure this to be enabled or disabled the Tracer then finally dumps the data it it calls stats analyzer and then it can dump the data into a file as a trace on file so how does that look this is sort of the output you can see initial block number current block number error per item confidence and these are the stats for example you have the pattern you have the you know the bites that it is and then you have the count that you've observed and you can go on you can specify how many you want to give an idea of the configuration that you can see all these are like this is the config that you can you can do you have enabled the file to write to the write frequencies like how many blocks you want to write this to ignore set like you want to ignore jump destination that's not really useful for uh for analysis so you can actually write ignore set you have an instruction Q size the size of the que used to gather instructions per block right because that's again it's it can increase as time goes so you want that to be configurable you have the processing Q size how many blocks uh are now you know are stored for processing uh you have the number of buckets that you can put in the sketch you have the number of hash functions that you can put in the sketch you can put the max error that you're willing to tolerate on the CM sketch you have the sketch minimum confidence that you want to tolerate you can put the analyzer top engrs to track like 10 10,000 100,000 you can put uh analyzer Min Supple threshold this says uh this is sort of like a filter uh both of these analyzer capacity and threshold then you have the sketch buffer size and uh you have the sketch re reset and reuse threshold which is like where it gets provisioned a new sketch for uh based on error uh so that's the plug-in config the next next section was patent Discovery selection and implementation so uh right so wait right so I used the stat analyzer I did two sets one was 10 U Top 10 patterns of two size two and that got merged then I did 11 pattern of 5 to eight op Cod that's still under review but these is basically in the patent matching mode of ilm so the the op code implementation has to be uh in parity with the evm implementation but you have certain opportunities of optimizing because you have a few patterns together but the major optimization will come in the Isle implementation because we have two implementations to do one is the patent matching and the Isle implementations the last was testing that was just started like a week week back so it's still like very much a work in progress uh so this involved testing IL testing um the patent implementations also the aisle implementations and the way we were doing this was basically we have two chains and we have an enhanced chain where we enable the ilm and a normal chain but we don't and then we compare the state route right but doing that also there are uh caveats because what if there's an out of gas error well both those Chains would be the same same what if there's a you know spec is not enabled again you will not you'll get the state route passing the comparison in valid jump destination so all these uh situations were there where it would just pass the test but it's like implementation is incorrect so it's quite hard to test this uh so anyway so there was some detection that happened over the weekend um I think five op codes uh that I fixed while testing and uh testing for the analyzer op codes and patterns are still you know work in progress and well uh thanks to all these u i mean my main mentors were Shimon and Iman uh and uh a big thank you to lukash Ben Adams Damian um for being you know sort of weekly there in the ilm calls and Marik and Thomas for actually enabling at the beginning to help me you know meet all my mentors and setting the meetings up so yeah thanks thank you so much lovely thank you
