Debugging data for Ethereum – ethdebug/format Overview and Project status | Devcon SEA
Devcon·Tue, Oct 7, 2025, 12:00 AM
Building debuggers for EVM languages is challenging, time-consuming, and brittle because compilers do not provide enough information to enable robust tooling. The **ethdebug format** project, sponsored by Solidity, seeks to address this concern by designing a standards-track collection of schemas for expressing high-level language semantics in connection with low-level machine code. Please attend this talk to learn about the status of this effort and a brief overview of its components. Speaker(s): g. nick // gnidan Skill level: Expert Track: Developer Experience Keywords: Developer Infrastructure, Tooling, Best Practices, debugging Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/
Transcript
[Music] it was about a year ago that I was at the solidity Summit in Istanbul and introducing the the work that was just getting started to build a debugging data format and so I figured it would be appropriate to come back a year later and report my progress and hopefully you know disseminate information about what I've been working on in hopes of getting people to to adopt the standard um I'm going to try and get to the interesting technical stuff but first I'll do some high level overview and that's we'll we'll probably run out of time but I'll make sure to leave room for questions uh yeah just just to get into kind of the the high level go over you know what what is this project what is it for why is it important uh making debuggers is hard and brittle and it takes a lot of guess work like for I don't know if you've ever looked at evm bite code before but it is entirely opaque and even if you were to peel back the layers of opak andess it would still be confusing and so uh I mean of course this isn't a problem just with ethereum and and smart contract Computing environments but also for traditional Computing environments but on traditional Computing what they do is they have the compilers provide more information and that has worked very well since the 80s we have debuggers debuggers even in the last 10 years are are really quite impressive these days uh but for smart contracts the the question is a little bit less or it's a little more non-trivial like what information do you actually need the compiler to output and how should that information be structured for this like our unique Computing Paradigm so let me introduce the project maybe not get ahead of myself so ebug uh the E debug format is an effort to build a debugging data format a debugging data format is the term that they use for the information that is output by a compiler that enables debuggers to exist at least to exist without spending way too much money building them so that's what the ebug format seeks to achieve uh we are funded by the EF right now we are uh proudly part of the Argo Collective spin out uh you should all go to stage two at 5:00 p.m. to hear more about that that's the last thing I'm going to shill apart from my own work uh just I get this lot I put on there we're not building at debug is not building a debugger we're building standards uh I yeah just we might build a debugger in the future but debuggers already exist I just want to make it a little bit easier and cheaper to build a debugger so who am I well I'm Nick nice to meet you all thank you for coming uh I'm the lead for the ebug project I am a member of the Argo Collective uh I have built a solidity debugger and then I uh oversaw it for 5 years and it is a nightmare uh I have worked on projects that are now ancient history and I just kind of work in the dev tooling space you know stuff interests me quite a bit so just to try and churn through these non-technical things project goals yeah we want to make a universal format I don't want a solidity buger only I want a solidity debugger that is also a Viper debugger that is also a feed debuger HG I you know and I don't just want to Target the use case of I am a software engineer that needs to figure out why my code went wrong I think it is also important that we target use cases such as why did someone else's code go wrong and lose a billion dollars so things like auditing tools analysis tools are very important and I would lump them into that same kind of category uh on that topic it's very important for ethereum and smart contract platforms that we target not just like debugging in development but as I'm sure you all can understand it is extremely important to be able to understand why a billion dollars disappears because of software so that's you a little bit more challenging than your your G debugger and and your your local build of your your C project or what have you um and because this stuff is hard it's clear to me that we we have to make it a goal to really optimize for adoption right I'm I have the solidity team in the room here supporting me in this talk and I feel very bad for the work that they will have to do to support a project like the eth debug format and on the debugger side it is only marginally easier but again it is orders of magnitude cheaper and simpler to take an approach like the one we want than debugging today ultimately right the idea is we want to lower the cost of understanding the blockchain right we want it to you look at the evm and understand what's going on or or have like a dozen tools to choose from that are all reliable and and are well architected right like this is a future that would be nice it's a future that or it's like a present and like your x86 architectures with your traditional languages and so I I hope that we can move in that direction for smart contracts so this is a a series of steps that I have followed many times and I would imagine that a significant number of the people in this room have also done this where they need to understand how solidity Works some like strange feature like how do mappings get stored in solidity or strings or or what have you and and you know many times the documentation is thorough and many times it is like even the solidity Engineers don't exactly know how modifiers work all the time stuff like that it's quite bizarre so what do you do if you're trying to do some blockchain like smart contract analysis well you sit down with solce and a bunch of input examples and you compare it to the output examples and you scratch your head and you think is this how solidity does it I think this is how solidity does it and then you go and you write your code and you you implement it and you you start like slurping blockchain transactions and figuring out what's going on and then you realize that oh actually in 0.8.6 there was a change but turns out I was just wrong actually also and then you go back and you spend another few weeks trying to make sense of how solidity behaves and how your implementation should behave and you know this is kind of annoying because people do these steps for the exact same problem like finding strings in storage like many people have done that same thing and why why why do you have to repeat the work so also that's just sidity people don't even try to the same degree with Viper like like yeah there are some Viper tools but like people spend a lot of time like I would argue wasting time making sense of solidity and there just isn't that for Viper but you know Viper also is responsible for billions of dollars of assets so I I think this is a a pretty significant concern that we we cannot ignore uh here's an example this one's at least documented uh like the the solidity docs are quite clear on how this works but it's still very bizarre yeah uh solidity has two different ways of storing strings if the string fits in 31 bytes or less it goes into one word if it goes if it's longer than 31 bytes It Go goes into at least one word and then you know at least two words and there those words are not consecutive right so like in the the first example it's what do I have slot zero and then the second example is slot one followed by a a catch a cash and then you start counting you know incrementing at using the catch cach this is very weird uh how does it work well that last bite at the end tells us two things it tells us the length of the string and whether or not that length is longer than 31 bytes and well the the second point is whether or not it's odd or even I mean I don't necessarily need to go into the examples you can look this up in the The Docks or you can see the gist uh but it's weird this is and this is not the only weird thing that evm languages do so how do you accommodate that when you know the the last you know 30 40 Years of debugger development has assumed that people are organizing data structures with like normal ways of organizing data structur well we don't have that privilege so here we are shall we get into it uh here's kind of the model that we've been working with right is like there there's there's e debug format and there are compilers and debuggers and there's some interaction between the three and uh I won't assume that the people in this room all know how compilers work so you can imagine well there's a source code that goes into the compiler in a highle language and there is machine code that comes out of the compiler in a low-l language and in between there is a long pipeline of many different steps uh some very complex some straightforward uh the compiler does things like when you make a function call and the compiler has to convert that function call into machine code it well it generates at least two jumps you know one jump to enter the function one jump to leave uh or if you are storing a struct or an array the compiler has to take your single assignment statement and converted into a series of like individual word assignments uh and in order to you know what you're left with if you're looking at the evm like the raw running evm is you don't necessarily see any way to translate back right who knows it's you just moved a word I don't know if it's part of a struct assignment or what uh but if the compiler were able to keep track of every single transformation as it performs it right oh I I have to convert a function call into a series of jumps well if you annotate those jumps to say this was the start of the function call this is the return of the function well and manage if the compiler can manage to preserve that information all the way to the the end of the compilation Pipeline and produce like a nice Json object well then the debugger debugger can read that information observe the running uh state of the evm and basically see oh we're executing this instruction what has the compiler told me about this instruction oh it is a jump that is part of a function call so with that debuggers can actually make sense like a coherent mental model of like the this like fictional highle world that is lost once the compiler is done uh and the idea is hopefully this should be sufficient for debugging missing billions of dollars in optimized code uh we do have challenges uh the the first significant challenge that has been on mind is the fact that when you debug something you are debugging runtime and unfortunately the compiler is not capable of predicting the future and knowing what runtime will be so like if you want to allocate an array in memory you have to put it in memory where does it go the compiler cannot guess what address in memory that array will have at compile time because maybe you are allocating five arrays or n arrays so we are limited in that the compiler can only know so much and there these gaps that like we have to take these like runtime observations like okay how do I take the the the running evm pair it with this compile time information and produce a a cogent model of the highle world that one's not too bad maybe of course optimizers make it even more complicated right when you have techniques like bik code D duplication right I mean optimizers are very important for smart contracts because you have to pay for every operation and it is significantly more expensive than than most computers uh so you know you might look if you're writing a compiler you might be tempted to say well oh this this series of of source instructions like this source code is very similar in its output to this other part of the source code so maybe I can reuse the bite code for those two same pieces like the two different Source ranges that like you know maybe they're in two different files maybe they're in two different functions whatever but they would produce the same exact bite code and the compiler would jump into that bite code in either situation so how do you annotate the this like lowlevel information suddenly you're going from well this instruction didn't necessarily just come from here it came from either here or here so we are effectively stuck with this like disambiguation situation where we a debugger will not be able to necessarily understand with full Precision like what what the debug information is saying but the goal is to uh ensure at least accuracy with that like maybe we can't be fully precise but we we can at least Target not being wrong and so it's like hopefully if you have a block of bite code that corresponds to two different originating sources you the debugger will have other information in its internal state to be able to say oh okay well I know that we were in this state so we must be following this code path and then another time we are in this state we must be following the other code path uh what do we have so far uh well we have a bunch of Json schemas I would estimate probably about 60% of our like working data model is implemented in schemas and the 40% is is a lot left but at least we you know have been working pretty hard on it like we have examples for every schema those examples are all tested you know I don't want to have to fix things that I didn't realize were broken in the future right um and on the op like the adoption concern right I do want people to use this format I'm glad you're all in this room listening to me because maybe I will convince you to to go and implement this uh if you do there are explainer document like uh documentation sets for all of the schemas and we have right now at least one reference implementation for what is the most uh tricky schema so far but the idea is that I think it's very important that we build reference implementations for all of the schemas hopefully on both the debugger side and the compiler side so if you're maybe you're inventing a new language or something and you can just you go to the ebug website and and see okay this is how an example compiler might implement this stuff I will copy that or I'm building a debugger I go on the website I see oh this is how a debugger might implement this technique and you know I really just want to make this as easy as possible these are the the schemas that are furthest along um oh I'm running out of time all right well I'll have to run through it so we have a program schema which represents a single piece of bite code and it's structured as uh a list of instructions with annotations for each instruction uh we have a type schema for describing variable types hopefully that's self-explanatory and then uh the most complex / complete schema so far is uh the pointer schema which is to describe data allocation strategies uh from a high level language and I'll go into a bit more detail on those so the program schema here if you want the direct link to the the explainer documentation for that uh this this scheme is quite incomplete but uh the foundation is there already uh the key concepts of the program schema are it's one program is one bite code programs have a list of instructions or instruction annotations those annotations allow the debug like describe a high level context and allow the debugger to basically have a lookup table when it reaches an instruction in the bite code it can go and consult the debug data and see what is happening with that um ultimately informing the high level State this is uh an example if you want to see so it's like the offsets the program counter I think with eof they're not really calling it program counter or it means something different so I just called it offset you can see the operation that that instruction is doing you might recognize this as the start of every single solidity contract uh in this in this example's case I also added uh a source range so you can kind of see oh this is where this is the corresponding line of code and uh it looks like in this example we already have a storage variable available to us so that's I got the first instruction we can read X which is of type string and it's in storage slot zero I'll get into the type and pointer schemas on two slides probably first I have to give you this warning yeah we haven't tested a lot of our assumptions here just full disclosure but we we think it's viable but stay tuned you might see me very happy or sad next evcon all right the type scheme is pretty straightforward uh you know you have to just describe types so we got known and unknown types like there's like uint strings like all the normal ones are supported by the format but we also allow compilers to Define custom types if they have you know some Advanced type systems uh from there types are divided into either Elementary or complex uh we do need algebraic types but they don't exist yet in this schema uh Elementary types are just like a single thing and complex types contain at least one other type uh types you know to D duplicate representations like there's a lot of data that we're talking about uh we allow types to be referenced by ID if you don't want to just copy long descriptions of arrays or structs over and over you can just give them an ID uh and of course if it's a user defined type it it can point back to the originating Source here's a couple examples it's pretty straightforward uh on the left is an elementary type for a fix Point number and then we have an array of ents as a complex type there's more examples on the the website if you want to find them um but this is really what I want to talk about is the pointer schema because this is uh I don't know I think quite interesting and suggests that this approach might be viable so we need to represent where variables live oh I should probably get this over with yeah so yeah I mean you guys can probably just go unless you like this kind of insanity then please stay uh but yeah so we kind of put uh Lambda calculus in Json schema and that's because of the compile time constraint uh which is you know the compiler doesn't know where something's going to be before it exists um right so uh basically you we get uh an expression syntax with this schema to describe like where is that string storage how do I know based on the state of the machine whether or not the string is short or long let me show you an example see oh yeah here here's the example actually yeah uh don't look at the slide just grab the QR code I I did document this there are comments there uh but you can kind of see at the very top this this string storage uh Al this this pointer starts with a storage slot slot zero and then we have a group of other data regions and variables I mean I'm really running out of time maybe I'll just I'm supposed to ask for questions now uh but you can you can see is there a Las a pointer on this well I could just point I really can't uh so so you have you have groups of basically the way the the pointer schema works is you have regions and collections regions are a single like continuous range of bites and then collections are this like abstract thing you can have groups of regions you can name the regions so you can have conditionals so you can do things like check the remainder is it odd is it even if it's even we know it's a short string so do this otherwise we know it's a long string so do that remember the example from before great all right I swear it's less bad than you think I mean the idea is that these are actually static representations so the compiler will only need to Output this once fortunately you know and then instead of zero it would just be a template variable and then the compiler can say use this pointer with this variable in place of uh string storage contract variable slot and this way way you know any supporting compiler can just output like have a library of these template these pointer templates and just output them whenever they're necessary uh and if they're not necessary it can omit them uh yeah apprehensive but we tested it end to end I like I put a string into a solidity contract in an automated test and I step through the evm and I observe after every string assignment that the value changes and matches the expectations so if you want to see how it works you can look at the uh the implementation guide hopefully it's coherent all right anyway there's a lot of work to be done um feeling very good about it you know I think I think this is a solvable problem if you want to learn more please join our Matrix chat we meet every other week you're welcome to join the calls uh you can also just watch the repo we got questions yeah thank you uh exclude the optimizer yeah so that's the approach we're taking uh will these get recorded I should probably repeat them uh the question is can we exclude the optimizer and just focus the debugger at compiler level Optimizer should not be altering code functionality unfortunately I think it does alter the code functionality but the approach that we've been taking so far is to start with just unoptimized code uh while thinking about you know how what do we do about the optimizer and making sure that we're not painting ourselves into a corner uh when we do want to support optimize code when is ebug format done I after Devcon others this is so impersonal biggest blocker uh I mean the the the tricky thing with some with work like this is that it's periods of challenging like thinking figuring out like how can I make this meet the requirements followed by months of tediously documenting and so it's like there could be blockers in you know either of those and and the different in both situations the blockers are are quite different uh can the debugger identify the function arguments if the format is followed yes that is the idea right so uh if I didn't have only a 25-minute talk I would have bombarded you all with way more examples they are on the website you you can kind of see um if you look at the program schema I actually handcrafted a pretend highle language source and the corresponding op code output so you you can kind of take a look at at what that looks like um how do I scroll down is there looks like there's more uh any plans on decomil decompiling deployed by code with no original source I mean this is a cool thing it's not really what I've my area uh I know there are decompilers that exist I I just haven't you know thought about that too much cool any more well thank you for your time everybody
Automatic transcript — names and jargon may be misspelled.