New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

CuEVM: GPU-Accelerated EVM for Security and Beyond by Minh Ho | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

Speaker

In this talk, we present CuEVM, an EVM executor implemented in CUDA for running a massive number of transactions in parallel. Its primary application is to accelerate fuzzing by testing transactions in multiple sandbox EVMs on GPUs. Additionally, we have integrated it into Goevmlab to support a broader range of use cases. We will discuss the design choices, challenges, results, and future plans to leverage CuEVM beyond fuzzing. Speaker(s): Minh Ho Skill level: Intermediate Track: Security Keywords: Scalability, Security, Fuzzing, EVM, parallel Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

[Music] morning everyone okay okay it's my pleasure to start the session today and yeah my name is min and I'm from the National University of Singapore and I'm working on parallelizing on the stuff and today I will introduce about our work on parallelizing uh evm execution so uh yeah let's uh go into the details so this is a very high level overview of what this talk is about basically we have gbus are very powerful accelerators invented decades ago mainly for gaming but then it founds it use cases in other uh domains as well people have used it for general purpose Computing uh scientific Computing and recently AI pick up the trend that's one make Nvidia become the largest company in the world by market cap so yeah on our um blockchain domain we found it use case in uh Mining and you have seen like people use it for mining ethereum for a long time before the merge but then after the mar a lot of comute power on on all of those uh sh views like they some of them come back to become AI training inference but some of them finds that you new use case in GK so you have seen quite a few talk in this Depon about accelerating Pro generating Pro faster using GPU but this talk is about something different so we want to find okay those gbus uh accelerator power what can we use what else can we use is for to help the etherum ecosystem so we think about one more use case to um Beyond ZK to run the transactions in parallel so nowadays most of the transaction you can see is not just simple transfer from one account to another so the transactions now are more complex smart contract executions and those are quite Compu intensive sometime that is also the one of the reason for people to set up like block limit and block scale limit so that the computation is bounded so you can uh verify run all the transaction in Block time and um so uh but we we don't intend to use uh this work to like run your transaction in Manet we found another use case in fing to secure smashr contract the use case we found from our work when we built our in-house fing tool to uh find buck in smart contract we got the chance to uh run and test and also compare with states of the art work in Academia and also in industry and we found that there's a very high demand to run transaction and test and you get some feedback you get some uh result so you can finally the main goal is to find the Box in smart contra to secure your uh application and GPU is just the main like just the the the the best like uh accelerator for you to run this kind of thing in this use case because it has uh the massive parallel programming model you can run like thousands of threats and how about we map those thousands of threads and thousands of evm instances running parallel so running transactions in parallel first we um focus on the niche fing use case but then along the way when we Implement our project we were approached by different teams uh startups they own building cool stuff and some of them use some of other uh sore to F out and have their own private Trel with their own amazing use case in the ecosystem but this m this talk will be mainly about fing I will have one slide about the other use case in the end so what is fing just a very simple uh high level understanding uh so first thing the main purpose is to find uh bucks in software it's one of the popular technique in software engineering and recently there's a trend for people to apply it to find bucks in your smart contracts as well so this is a very a high level depiction of graybox fing basically you have a tool that generate a lot of inputs those inputs have more coverage and more complex than your unit test usually you use so it's more complex than they eventually it's a iterative uh process you generate input you execute it and then you get feedback so those feedback can help the furer to like generate better input eventually you can find the input that can trigger the bug in your smart contract or in other software so the input are the so if you map it into um ethereum transaction to test Mar contracts the inputs are the state and transaction to interact with your D your smart contract and then the feedback can be traces branching data or the updated state after you run the transaction and those are up to the uh fing tool design and implementation and yeah so one example on the right is how they instrument the evm to catch box I just so very simple example some of them are not uh like the add overflow is not very uh like it not exist anymore after solidity eight because they inject by C to check it before they try to revert it before the Overflow happen so for example if you want to catch overflow you just instrument the event to every time you see upcut ad you check the operand if it's like over uh the the two to the^ 256 then it's overflow yeah or another interesting upcut is jum ey on of your if else condition in your smart contract or require statement it will be encoded into this jump eye and those will be the helpful hints for the fing tool to further generate better input I'll show one example of the Run uh after after this so this is first thing on a conventional iterative process on your CBU then we try to map it into GBU execution so you see um our verser can generate thousands of input instead of iteratively one by one now you can generate a a bulk of them like a batch of thousands of them and then you can offload the execution of thousands of them into GPU so this is the main uh idea the main uh proposal and that that uh our project trying to achieve so so after we offload the execution to GBU we can map one or few threats in GBU to run this uh transaction simulate the evm and then you get the feedback back to the verser on the CBU and try to update it internal stage strategy generate better input and yeah that's about it you have to uh work you have to uh have some Comm communication between uh the accelerator and the CBU side and they all work together to find bucks in your smart contract okay so we offload the evm execution to uh GBU and AVM execution basically you execute one uh transaction and then you given the world state so this is a very high level of what is running inside so usually nowadays like uh transaction interacting with smart contracts are quite complex they have Dynamic Behavior it's not just one single evm call you run and you exit it's actually uh every time you have a con context and then each of them have the volatile machine States like stack memory and they have their own size and it's like evolve over like the running at the run time we don't know beforehand so this is also one of the challenge when we try to implement on GPU we canot just preallocate the for example you know the stack size uh maximum stack size and maximum call dep if you preallocate everything is is gigabyte so and we we don't need to do that uh at the beginning so that is one of Challenge and when you have call context and then you can create a new con context or you you can revert back or exit and update the parent State uh depending on the result of the the child call so these are the high level and when we map to the GBU threats actually we only consider each of the the current call context uh so when you enter a call context you have your own St memory uh volatile machine States and then it will enter the execution Loop basically you just keep fetching uh new up Cod and increase program counter executing the logic and then um eventually this context either uh return back reverse exception or create a new C context and it keep looping um and then we we have the decision of how many uh GBU threats that we can map to execute one uh EV instance and imagine we have to run thousands of them so the decision is actually uh um dependent on one of the library we use uh we use a ccbn library from Nvidia lab to reduce the complexity of development at the moment so they have the limitation like how many threads we can configure to compute one you into physic um arithmetic operation so imagine it's like some big in uh library that you can use out there and um yeah and also because they have some limitations so we are currently have a plan to develop our own uh library for this they are developed for uh very general kind of use cases you can configure the bitwidth but then for our use case you only need to use 256 Uh u in so yeah we can optimize a bit more on this so this is how you map the evm to sh threat and um after we Implement Ming using those library and you need to ensure it's correct and uh for this one uh when we develop this one we didn't know much about the tools in the ecosystem so thank for the advice from etherum Foundation we we actually need to implement and compare with other EV as well so uh we um we integrate with go evm lab which is the uh the tool developed by core developers of ethereum and then it's available there and we uh make sure our trades are compatible with other event as well that we com compare line by line so we use eiib uh 3155 make sure it's our both correct and then um we compare and then we uh use a test to compare and then uh currently we pass on the test and functional folders that the important one and the one I highlighted here are the one that are most time consuming to us or recombine contracts and zero knowledge they are just a few recombine contract but the logic and the time we spend on them is quite a lot so yeah uh evm of course is simpler than this and some of the things we can use open source but some of them we have to develop on our own especially uh easy pairing actually we copy from the uh python library from ethereum so we write it in Cuda based on the logic there and um so after the all the test so rare Corner cases remain and mainly gas difference uh so yeah so if your uh F this case is not Reliant the logic not reliant on gas termination or some of the the logic then you can actually use it correctly so we are still trying to uh fix on the remaining gas difference so after we are satisfied with the correctness then we think about optimization uh so because time is limited our developing time also the time of this talk so I just uh big one of the optimization that we did recently to show what what we are doing so basically um some of the we we try to optimize the normal uh transaction that you usually uh see out there in the real block and the real Transaction what are the popular up course what are the things that critical to Performance what are the things that you have to uh execute all the time so we optimize them first so basically we check the statistics and then we found okay it's it's quite intuitive it's quite obvious that the stack operations is almost everywhere you have to optimize the stack the first and then after that we also check the the stats of the stack size what is a normal stack size that usually transaction use and what's the memory size uh then the program counter it go up to 5 8,000 so we found out okay we cannot cat the program counter it's too much sorry cat The Bu Cod 8,000 bite is too much but then stack size 20 something we can allocate the fixed size uh stack on Fast memory that's what we did so we should alloc pre-allocate that one first very fast uh uh stack on sham M and then after if you your stack is bigger than that you can use a slower memory as well so most of the transaction you use this you don't use a lot uh you don't use large uh stack then it will run very fast H that's the one of the main uh optimization that we're trying to do we try to optimize for conventional common transaction you can find out there okay so current implementation we are happy with one uh Beta release currently so basically uh we support Shanghai we started with Shanghai but we couldn't catch up with new uh evm hard for it's quite a lot of e implementations so we stop at Shanghai first and then uh executable we output the Jon trades and then we have two versions CBU GBU and we also have the share Library so you can use it in two mod one is you use the this share lab Library it's open source now and then you can use it to try on Google collaborate uh collaboration the Google collab you can use uh GBU there so uh yeah you can use in two mods you can write your python code to interact with it and also you can use the executable um yeah so performance is still depending on the workload um so it requires more comprehensive benchmarking but yeah we compare with our own CPU version is faster but one one thing is that uh our CPU version is actually slower than the current um evm implementation out there so it's not like directly Appo to Appo comparison and yeah we are still improving the performance um yeah so this is one of the sample run in fing and real uh like in action so you can try on Google collab the link in our GitHub Trel and then uh one example is a to example in smart contracts with some of the very obvious Buck there and some of the input guard like if conditions so to make sure that you do some real vering you make sure you get the correct input you can run it and eventually it will like output the lines the S line and the input that Trigg go the the bug like you see input equal uh 400,000 or some uh yeah uh 40,000 so inut 4 five 6 7 and Trigger the bug it can show you the S SC the line of B this one running the execution running in GBU so so besides detecting bug we also have feedback so this are feedback ranging feedback one example is this function test Branch uh it's like there's a condition for input to be uh 1 million and you see the current input and the distance before you can break that uh Branch those are the gray box F technique that we Implement in here those are quite common in uh great block fer uh uh fing tools out there and our current relay the botton neck is still remains in the way we uh repair transaction and we get back the serialized data to update the state of our fing tool so it's stillo slow but we have experiment experimental Branch where we remove on of the complex logic and of course this Branch does not confirm with the yellow paper anymore but then it will reach very high throughput it can reach like 60k uh transaction per second and improving and we saw some teams uh have private trable and some of them clone or for from our Rebel with with like private implementation on optimization they can reach th 100,000 of transaction per seconds as well but we are open source and we're trying to make sure it's still confirmed with the etherum standard and and uh Beyond fing this is uh some of the interesting use case where we talk with other teams uh other team doing Co stuff out there but they are doing the layer two and also fing uh but the main thing when we we were thinking about baliz transactions for ethereum execution client is it possible is sound cool but then it's not very practical at the moment because you need more transaction currently how many like hundreds you need uh you need thousands to have like to to um exploit that bism in GBU and also the transaction they are all different they they are not very uh similar so it's not easy to uh achieve speed up on GBU and also to achieve this kind of thing you need more client support like the GBU memory you cannot just get that terabyte of War States is not possible it's in the minute it's too big to run and also because of memory concent and also uh the client need to ensure it's safe and correct to run those transaction in parallel which is not because transaction you need to run sequentially to one by one after the other and eventually to get the final State this one run in parallel there you have to make sure that all the N running it reach the same deterministic out put so they can reach consensus otherwise it's it's not possible at all so that's what we think at the moment so we still can find other use cases be besides like trying to uh speed up ethereum execution client uh we see other layer two teams also can try to have their own uh execution uh engine also transaction simulation platform I saw like people use it to simulate swap for example and imagine you have a lot of similar transaction testing simulations and this one is a use Cas for it yeah you can simulate million of squ in a short time serving a lot of users okay so uh the last two slides I want to talk about our team and collaborators so yeah working with me uh also uh we we have a small team of three uh researchers and Engineers uh working with me is ch leave and we are from the Singapore blockchain Innovation programs uh n us also working with me was Dan and uh we work in Singapore before and he went back to Romania to become professor and uh yeah we want to thank Fredick and etherum foundation for the advises and the grand support especially the advises to make sure it's correct and how how to uh use all of the new tools in the ecosystem that we were not really familiar with thank you yeah so that that's that's all about this one and you can find out on the GitHub and um yeah just feel free to create BR open issues discussions and reach out to me uh I will try to help thank you man and thank you for this amazing presentation um so I have a few questions let's go through them let's take the first one how much work would it take to integrate this in an open source fer like a kidna uh yeah you you can use it currently we release the the uh the share Library you can use it right way but the thing is you need to modify your fer our our example toy fer we have that bottle neck to repair thousands of transaction before you send to GPU that one is quite a big Buton neck you run it on GPU is fast but then you prepare for it and you run you collect the result is still quite bad so currently you we need time to like improve the interaction between the library and your fing tool so it works but you need to work on piping yeah the interaction yeah okay on Plumbing all right can qvm plug into sorry qvm plug into a simulated node tool like Anvil or E tester yeah so in the end we need to make sure is the it's possible but we need to make sure it's compatible for API call so basically we provide that Library if you have your own adapter your requirements so you just send us the state and the transaction you want to run but should make sure the state is not so big we can manageable like transfer it to CBU and keep it there sorry GBU memory and keep it there then it's possible to transfer the state transaction run and return it back to result uh whatever result forat that you want fantastic thank you all right third question would fuzzers after change to adapt your GPU acceleration it seems like this is a different version of the same question let's mark it as answered all right fourth question how does the GPU Fred get the state from when executing s load well that's specific okay you want to take the next one so we have the data structure and reside on global memory so for the stage me and memory we have to keep it in the slow memory we don't do any caching and that one we keep the structure pointing to the state and we set through that so basically if the state is Big so this is one implementation detail we didn't Implement very fast searching in the states basically we have to go through on the array of accounts so if the state is so big currently it's quite slow but yeah uh that's how it works so everything on global memory we have to set through them fantastic thank you all right we still have 10 second 10 seconds so maybe a more playful question have you considered calling it cute VM rather thanm Q VM I but you know suggestions you know yeah yeah we might change it to this team wonderful I mean thank you for your time and all the work you obviously put in your uh in your slide and um looking forward to your work yeah thank you all right

Automatic transcript — names and jargon may be misspelled.