New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Protec and Attac: Programmatic Execution Layer Consensus Tests by danceratopz | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

We'll give an overview of Ethereum Execution Spec Tests (EEST), the new Python framework used since Shanghai to generate test vectors for Ethereum Virtual Machine (EVM) implementations. By generating tests programmatically this modular framework allows test cases to be readily parametrized and dynamically executed against clients on live networks. It tightly integrates with the Ethereum Execution Layer Specification (EELS) and could potentially be used across the L2 EVM ecosystem. Speaker(s): danceratopz Skill level: Intermediate Track: Core Protocol Keywords: Core Protocol, EVM-equivalent, Testing, pytest Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

[Music] all right yeah um thanks a lot nixo for the introduction yeah my name is Stan ceratops I'm a protocol tester in the ethereum foundation and today we're going to talk about a relatively new uh repository and test framework for the execution layer but why do we need a testing framework anyway I mean like how hard can it be right so the main well client diversity is like a foundation of ethereum core protocol development and as you heard in ethereum's talk in the opening ceremony that having two clan implementations for the 2016 Shanghai dos attacks really helped the network stay online um during these attacks so client diversity having two clients with two different implementations different code paths can make the client more resilient against attacks or software bugs and today we have many more implementations so we've got at least nine evm imp implementations across eight different languages and I'm not even talking about exotic implementations I'm talking about things that I generally work with day-to-day but having so many implementations does add quite a bit of complexity to testing and so we exist as a team East and it's also the name of our repo ethereum execution spec test which and we maintain a set of client independent tests um that every client team can use to verify their implementation of the execution layer so um what are we testing and how do these tests look like so the main thing we're testing is that no matter what state change we provoke in clients by by sending transactions to the network that afterwards every client would produce the same postate so the ethereum world view is exactly the same amongst all the clients so we have two main formats in order to um or test formats in order to test this we have state tests and we have blockchain tests and the idea is more or less the same that we have a pre-state which is like kind of a Genesis allocation for the for the for the blockchain we have a collection of EA accounts or smart contracts that have been deployed on the Chain that's the pre-state the setup for the test then we execute a transaction in a specific environment for the evm and so what we do is we excite the evm with a given pre-state and environment which can be for example like block specific parameters by sending it transactions and then we see um what the what the generated postate is and we want this postate to be the same across all clients and blockchain tests are more or less the same but they test um so if you think of a state test it's like more granular so it's on a finer scale you can test individual transactions a blockchain test can not only have individual transactions but also um block contain blocks and these blocks contain contain one or more transactions so they test slightly different things or concentrate on different things so State tests can test op codes or gas usage and a main part of our work is really to verify smart contract interactions um and we can verify that trans transactions that should get rejected do get rejected and blockchain tests validate block progression from block to block block rejections and Fork transitions so before we go to the new let's have a look at the status qu at the merge and see what the state of things were for testing at the El um back when we met Dev in bogot and Devcon 6 so since Frontier there's the repository that's manag test generation or the body of tests is ethereum tests and this is a huge Corpus of tests mainly created by Dimitri who's sitting right over there um who did created an amazing body of tests um later joined by Ai and you can see that there's almost 5,000 tests for Cancun and this these test cases are written in yaml and typically by hand um and they can contain they contain comments as well about what the test does and then they get generated by uh a tool in another repository called retest eth which generates Json test fixtures using a reference evm implementation so the test fixtures are then so this process is called filling the test and the reason is that the test fixtures contain um much more information about the test in particular all the changes in the post in the state and importantly the state route which basically the fingerprint of a change that you can really verify that every client did exactly the same so in 2021 AI had the idea to create more complicated tests and so instead of just writing these yaml files by hand he started generating uh many yaml files using a node node script like using JavaScript and this allowed him to create many test cases um by by looping over op codes for example so we get to programmatic tests and generating tests programmatically so around the same time in 2021 like client was at one of the first interop meetings in Greece um was working on account abstraction I'm wondering how can we test that there's so many parameters that we need to change in so many test cases that we need to write how can we possibly manage this complexity by writing yaml files and he also came to the same conclusion that we should have programmatic test generation so what like Cent did was um he sat down after the inup and wrote a testing framework in a repository called testing tools which specified the tests as Python and then generated Json from the python and so this makes it easier to Define new test cases and parameterized test cases so he he worked on this for about two months and then this repository kind of went into hibernation and didn't see much work on it well actually no no work until Devcon 6 so at Devcon 6 everyone was ecstatic that the merge had gone so well and it was very forward-looking that people were seeing what how we going to handle testing all these complicated features that are coming into the evm in the upcoming hard fors and so there's a discussion about what what kind of framework can we what can we use or what should we write in order to handle this and like client piped up and said I I think I've got a project that that could be useful and so this became ethereum execution spec tests which is the repository that I work on today so Mario sat who was a protocol tester at the time sat down and um started writing the first test cases building out the test test framework a little bit broader and was then late like soon joined by uh Spencer in January of 2023 and they wrote tests for all the eips included in the Shanghai hard fork and in the first release for Shanghai or the final release before the fork you can say they had 194 test cases to handle the new features coming into Shanghai so Shanghai really showed us how powerful a programmatic test uh test framework can be and how easy it is to write test cases in Python at that time the repository used completely custom code to collect all the test cases and like generate the test Json and although we're not writing test like unit tests we had the idea what if we just use a standard uh uh test framework to do the job for us and so we rewrote the framework to use py test um which is like the de facto um test framework for Python and pest has a lot of advantages it's extendable via plugins and there's a large existing ecosystem for plugins so if we ever need a new test result format we can just install a new plug-in and get the test results in Json um or an XML or an HTML um you can also add your own custom uh plugins which is very important important to us because that's basically how we've implemented all our commands with in East in order to generate tests and later we'll see execute them it allows for easier debugging and last but not least it means that we can really easily parameterize test cases and as you can see in the screen chart which is from our documentation um how heavily we parameterized um um blob transactions for example so what you're looking at is test valid blob transaction combinations and so this is a blockchain test that um has one to six transactions um and each transaction contains a different number of blobs and we test all possible combinations that they're accepted by clients and we have an accompanying test for invalid blob transaction as well and so essentially to parameterize this test becomes one one liner and if a spec would change and the max blob uh count increases we can change it in one place and generate many more nice tests so this is really accelerated test development and you can see in the Cancun release we had over 2,400 tests including that just for Cancun so this has also really helped us to create a contributor first repository we really want this should not be a dark corner of protocol development we really want EIP offers EIP implementers even white hat hackers anyone to come along and use this repository we want them to come and write tests there going to soon going to be a new tool to help you get started interactively thanks to Rahul who was being an EPF fellow and we'd also like people to be able to generate tests from a transaction hash of a bad transaction on a devet for example um automatically so you generate the python code directly and we also want people that even if they're not directly contributing to the repository EIP implementers after they've been working on their implementation and wondering if Corner cases are covered in the test cases they can come to our repository and look at our documentation browse it and really check whether the coverage is there that they expect is the corner case they just thought of is it already in the tests or should they come and pest us to make it or even add it themselves this has really helped us unlock parallel like development for multiple Forks in parallel um and for Prague we've got 2,500 test test cases and in parallel we've been also implementing tests for the upcoming fork for the evm object format where we added a new test format to validate validate eoff uh B code and we have nine 9,000 new test cases already for eof and this is really a huge thanks to the ipsilon team and to Dano who have done an amazing job of um writing these tests and contributing to our Repository also a big thanks to ignasio and Gom and Spencer from the testing team um who have contributed to adding the infrastructure to the framework in order to test vericle so it's already possible now to fill and run existing tests until Shanghai for the vericle transition and the witness has already been added to the fixture format so we're ready to test that when we move closer to veral so at this point we had a relatively usable filling framework for generating these tests but we really missed something we really missed the feedback so if you want to really test a client you can't just fill the tests that's just generating the tests and we mentioned before that we use an evm reference implementation to fill out the features so we can get a state route that we can compare to other implementations this is a differential test and if you want to really test a client you've then got to take this Jason and then test the client by getting this other client to consume the Json that the first client produced so it's a two-step process and we were missing fast feedback within our repository so we then added three uh commands for um us as test developers to quickly run the tests against a client the first one consume direct you can execute or run the test against the C native consumer and the last two consume rlp and consume engine um allow you to test a client in the hive test environment using different C- paaths either by loading rlp encoded blocks upon startup or via engine new payload these commands that were in thank thanks to the modal AR architecture we've created using py test plugins and um we really use combinations of these plugins in order to achieve um the commands that or the the tooling that we we need in the reposter and Mario came up with an excellent idea that to extend the framework to allow execution of these tests on live networks so he added the execut a command which is allows us to attack live Dev Nets which we just did in the framework uh in the workshop a couple of hours ago so the execute command doesn't generate Json fixtures it takes the test cases generates the transaction and executes it against a client using the its rpcm point this really unlocks a totally new use case for our tests and really allows us to get out of a staging like setting that we've been executing these tests in until today and run them in a more production-like setting so we can really run these tests against Dev so we've done a lot of work over the last year or two years to improve the lives of test implementers but we're not done yet we'd really like to also help improve the lives of client developers one large job but still not done is we'd like to merge e the the Corpus of tests from ethereum tests into East and um we'd also like to um this would mean that clients can then they only need to download one artifact because ethereum tests are still they're still very important to consume and um at the moment client teams have to develop uh to consume tests from two different sources and once we manage to merge these tests we can catalog them a little bit a little bit better and clients only have to depend on one artifact to test their clients and as you can see we've developed quite a lot of tooling over the last couple of years and we'd like to be able to provide this tooling to client teams so they can use it in their test Flows In order to get faster feedback on on their performance on a PR basis so they can rapidly iterate on the development for that we also would would help if we had faster test execution Hive is quite slow for example because this test environment is a system test which instantiates a client in a docker and so and this this client gets the docker container gets pulled down and started up for every test we'd like to make this a bit faster so clients can get faster feedback additionally we'd like to enable uh easier test debugging so that clients can really run um the same test hopefully if possible in their native environment so they can drop into their debugger um when they need to in order to detect to find out what the problem is when they're running the test we also like to improve the development cycle so our team is recent the east team has recently merged with the execution specs team the eels team that handles the P for implementation of ethereum specs we we've merged with them to become one team called steel the specs and testing of the ethereum execution layer and we now def uh fill all our tests the reference implementation fulfilling test should really be eels and it is for all Forks up until including prag and soon prag this would unlock a kind of test driven development it's a rather coar test driven development because you need an endtoend test obviously for your client but it does mean that when a client imp implementor comes to write his Implement his code for the EIP he will have tests ready to go and he'll be able to check whether it's there so the F feedback will be fast instead of having to implement and perhaps wait for tests to arrive from another client another thing that we'd really like to add would be EIP versioning so there's many different places where an EIP is defined or implemented there's the ethereum E's repository there's um the test and then there's all the client versions and the specs of course so there's multiple versions where the eips are defined and if we can track them versions across all these components more easily we'll be able to make a statement about whether these whether these components are even compatible before we test runs uh run tests sorry so we can even just basically um xfil the tests and we know that we expect a failure and we don't have to try and debug something that doesn't make sense because the specs are just incompatible and last but not least we'd like to improve coverage it would be amazing to be able to hook a fuzzer into the tests so we already have a whole bunch of edge cases for eips and if we could hook like Define which parameter you'd like to fuzz for these edge cases um it would allow us to run these via execute and um get more test coverage and we'd like to improve our documentation even more and encourage people to come and look at our test cases and try and contribute to new IDs ideas to increase coverage so I hope you see that the title of talk is not just a meme um our repository does protect by allowing clients to consume the tests but now we can also attack clients on dev Nets so thank you very much um one of the main messages of the talk is really that we're here and any client ever anyone who's interested in testing on ethereum can reach out to us and you can find all our contact details in our user document or in our repo documentation just go to the docs but you'll find Linked In the repo go to getting started and getting help all right thank you very much let's take some questions starting with the most upvoted can you programmatically check test coverage that's an excellent question and uh Dimitri's been working very hard on this and uh we now have um a GitHub a GitHub action that actually does exactly this so it detects if um if a test has changed and then it will basically fill the test um before and after and run the fixtures against evm 1 from the ipsilon team and get a test coverage um on evm one before and after and then literally Check Line like what's the line difference in in coverage so yes were significant bugs found using this testing framework uh I mean we hope to catch the bugs early right so we we don't want to catch the bugs on devet really we want to catch the bugs as soon as possible while the client is implementing and so yes but hopefully you'll never hear about it up until Petra which EIP was the most challenging to test up until P not including Petra not including pectra hu let me think I would say probably 48 for four yeah are test fixtures in ethereum Spec tests already included in ethereum tests um yes but um we'd rather go the other way so um we consider the future of testing ethereum spec tests and so we'd rather get ethereum tests in execution spec tests as an interim like solution dimetri included or filled our fixtures from retest e using our framework and included them in retest e uh in theerum tests like the test repository but because our tooling is our repository and we can generate releases more quickly I would prefer to keep our releases in execution spec tests and I think that's the approach we're going for now and hopefully soon like before eof we'll have ethereum tests and execution spe tests can you give some insights into the testing limitation like what is not easy to test um I mean definitely one of the challenges is keeping up for specs so it's not easy to test a moving Target which is maybe not the answer you expect but um it's definitely very hard to to keep up with the pace of ethereum development and have tests ready for clients to go is the reference evm independent of other clients yes and I'll just say again the reference evm implementation should be eels that's exactly what it should be what does filling from eels mean filling from eels means that we have a test case which may just be a few lines of code um that describes the entire test case thanks to the framework and Eels and Eels basically um so the the way it's done is there's a tool called the trans transition tool otherwise known as t8n and this gives a doorway to the evm and we give it um inputs in a transaction and that and this evm the TN tool gives us the expected output and so this this means that we generate a lot more information as an example um the test framework can't tell you what the balance like we could add it but we don't tell you what the balance of an account is after it's sent transactions to a network because we don't know that we we could add all the gas costs but there's a lot of complicated things that we need to track and we don't track all of them so if we wanted to track all of them we'd be implementing an evm and we don't want to do that it's not our Focus so we use a reference in evm eels that gives us all this information including the state route will the engine API get incorporated into this version testing mechanism um it is partially um because we generate different test formats and one of the test formats is in uh is called the the blockchain blockchain engine test format and um these these fixtures can be sent to a fully instantiated client that's running within within a hive environment and they're executed using engine new payload but the engine API specifically this is really about more evm specs rather than engine API specs there's a separate simulator in the hive repository that handles these tests great thank you so much we are at out of time for questions but if you have one of these questions find Dan afterwards and we will be back in 5 minutes

Automatic transcript — names and jargon may be misspelled.