# zkProving the history of Ethereum in real time. by Jordi Baylina | Devcon SEA

- Speakers: [Jordi Baylina](https://streameth.org/speakers/jordi-baylina)
- Channel: [Devcon](https://streameth.org/devcon)
- Date: 2025-10-07
- Duration: 26:46
- Watch: https://streameth.org/watch/yt-boSCLHs30tk
- YouTube: https://www.youtube.com/watch?v=boSCLHs30tk

## Description

I'll explain the current work that we are doing in the Polygon zk teams to improve the performance of the provers and the quality of the tooling.
I'll will explain how we can parallelise the generation of the proof and how we can integrate with different hardware and software so that it should allow to build a zk proof of a block in real time. 
I'll explain also how this proofs can be recursively linked to build a zkProof that can proof the whole Ethereum history from the genesis.

Speaker(s): Jordi Baylina
Skill level: Expert
Track: Core Protocol
Keywords: ZK-EVMs, ZKP, Zero-Knowledge, lightclient, type1, starks

Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon
Learn more about devcon: https://www.devcon.org/
Learn more about ethereum: https://ethereum.org/ 

Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more.

Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. 
Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024.
Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

## Transcript

[Music] hello everybody um we need fast proving at the beginning was we need proof we need to generate proofs at some point we need to generate cheap proof this was the beginning of the zkm but what we are seeing now is that no proof are would say it's cheap enough especially if you are doing a rollup and you have a transaction but we need to proof them fast why we need fast proofs I would say for many reasons uh the the reasons I mean in polygon for example we need it very much for the ACT layer when you when you want to transfer funds from one chain to another chain U two things needs to happen first you need to one chain needs to commit to one state so that the other chain can use the fonts of the first chain and the other thing is that the second chain needs to be sure that this state is correct okay so it has two options the first thing is that the second chain just follows that chain and then warant is that this this this state is correct or the other option is that the chain a provides uh zero knowledge proof that uh his state is correct in the first case well it works well maybe if you have two networks but if you have hundreds or thousands of networks that are appearing or disappearing every time it does not scale well I mean you cannot ask all the all the all the all the validators of all the chains to run a validator of all the other chains that are in the aggregation space so the way to the way to uh U scale that is by having a zero knowledge proof that uh warrant is at state is valid this is um um and then but then this means that the user when he's doing some interchain communication uh in the user in the user interaction I mean in the user in the ux you will need this uh you need to Wi for this time okay so and uh so we need to generate these proofs as fast as you want other context I mean yesterday in uh in the presentation of The Bean chain was a ZK uh chain so that means that in in this context you need a lot of places like clients needs to validate so needs to prove that the the full execution chain it's another case that's important so in many cases this chain uh so generating fast proof is important so how fast can we go and I want to present here so an internal project that we have been working for uh the last five six months uh inside ethereum we call it internally zisk and is very much about that is how much we can push the proving system zis is also based U it's a system that can uh prove u i mean you can build a pro programming rust it's a r five uh compiler is a r 564 bits it's a little bit different from other projects but the full idea of this project is how fast can we uh can we move how short how fast can we generate this proof this actually are this three different I would say three different pieces or three different projects itself one is the pr itself Pro is a generic Pro so the idea is that you can build uh your own Pro that can proves anything which say It's Kind the hardware layer you you we have the pill two pill two is based on um pill two is based on um the pill one language that we use in the zbm and allows you to do an arithmetization and way we extended a lot of the a lot uh of the we extended a lot the pill to language the the pill to language uh I mean making them more easy for auditing uh this arithmetization and mainly including um one thing that's called vcops that mainly means that instead of having uh monolithic proof we can divide the proof in sub proofs this is very important because then we can paralyze very well so we have a proof we divide the work in different sub proofs and then we aggregate them together this is the main idea of how we can accelerate uh uh the things okay so the the pro I mean it's a generic Pro so how it would work okay this is the process that if you want to build a any proof it can be from a fibon serious if it's the kind of a Hello World until uh I mean a VM project that's in there or maybe some whatever you want to build mainly you need to build two things so you need to build armatization the aration is in this um in the in this uh so we just write the pill so is are in meditation itself you need to write what we call the windows computation this is at the end is filling it's a program that's actually filling um the trace of this and that's very much and with that you already have a so with this um well you compile the pill okay so you get a kind of a pill out format you get the rust you compile uh a library I mean you get a a dynamic Library okay from the pill you get some helpers so that allows you to build this Windows computation easily okay and then here we just run a normal setup Pro setup here the setup is the configuration of the pr itself some Starks configuration that's just a configuration file here we generate a normal approving key and a and a verification key and then we have the normal thing we have a approver for the approver have an interface that you get an input and this Pro generates the proof and the public uh uh and the public outputs and you can take the proof and the public outputs and verify and the verify you see that's got it here okay so this is the so this is would be the schema that would that would work okay um thing is that this Pro is generic I mean you don't need to recompile the pr it's just a u just a program that's running there but can be a program can be a service uh can be I mean um this PR and this program can have like different versions I mean it can be in GP GPU in I mean you can it doesn't change I mean you can have it's a generic for any the co the idea is that it's generic for any for any circuit this program is I mean is is U it's open source um and is uh everything is designed to minimize the latency okay we will see later uh the different things on related to that okay so and then uh um let me just go back so um this would be the prer okay but on top of the prer we have a specific circuit so it's a just just one uh so we are using it's like a so the specific circuit using p 21 that's implemented in P 21 we can use the pro that's actually this um uh this dis is this uh let's say RIS five emulator kind of actually is not I would not call it an emulator it's a ZK processor the thing the only thing is that uh there is a direct conversion between R 5 or even wasam or other or or llbm to this ZK processor it's a processor that has only about uh 50 polinomial so it's very very lightweight uh processor and really fast and and really fast processor okay so um and then it works very much like any other VM in there so you can build a program so you you build the the circuit in in a normal program of course you can test your circuit at the end it's a deterministic program so you compile the the the circuit you get a program and then you have an input and an output it's a deterministic program that generates the circuit once you have that that's the only thing you need you can take this program you can generate mean you can compile using rust and some setup and then you get a kind of a ROM or if you want a you know an elf kind of a program itself um and then uh you can use the Pro I mean the pro that we saw before of course here the the the zis the private key the the verification key is already is already done because this is a an a specific application a specific circuit for the prover but the prover is exactly the same and then of course you have an input uh a proof and output and the verifier the only thing is that the input includes the the private private input and and the ROM so you can prove any any program that you want and of course in the verifier you can you can use the ROM has so you you can use the fire itself okay so let's go deep in how uh this I mean how how the architecture I mean how this um uh circuit works well as I told you before this circuit uh is based on uh vops actually is many there are many sub proofs each this Su prooof there are different kinds and this this works very much like a normal processor we have like a a main processor with a program with a ROM and then we have a ram we have a arithmetic operations we have a lot of co-processors um uh that can be connected here okay uh so each one of these is one some proof and actually you can have for example arithmetic if you have a lot of arithmetic operations you can have a lot of sub circuits and all them add them to the bus and they connect them together this bus I mean is not an electronic bus but there something would say similar it's more an algebraic boss but this um we are using here lock ups and here the idea is you can have Su proof so you can have the main processor you have the arithmetic processor each one is like sending things to the BS and the main processor says okay I'm assuming that this arithmetic operation is correct okay and then there is the arithmetic coprocessor that's saying I'm guaranteeing that this arithmetic operation with these values is correct okay so the full system is correct so one is adding to the BS the other is substracting to the boss and at the end the boss means zero if if everything is uh uh warranted and the cool thing is that I mean in this boss I can add arithmetic operations I can add memory instructions I can add whatever you whatever you want we can have we have a kind of identifier but with the same polinomial we are just adding and substracting different things all all together okay you see the arithmetization I'm not going to go in detail on the arithmetization but I give you the the big idea I mean um the processor is not based in registers registers in ZK I are relatively expensive mainly because in one instruction you are using in general one or two registers so that means that there is like if you have 32 registers you may have like 30 30 columns that you are not even using in that line so doing registers in ZK is quite suboptimal so the main idea here is okay we have like so in each so in each instruction what we do is we go directly to memory we are doing two access to memory we are reading to things and we are storing back to the story back to the memory and in this and this so this is the the the main idea okay so we have only we have only one one register actually we have a couple of registers a program counter a kind of a frame pointer or stack pointer um in there we have this main so this main instruction that can be any instruction there so we have a way this is a way to load the first register load second registers I mean you can load from memory or you can load maybe from an immediate value or other um uh things maybe you can take just the prev see and things like that okay um we um we store the memory we gener we store to the memory we can also for example store to a program counter so we can do a dynamic jump in there okay and the we have the operations the operations at the end is just throw into the boss the operation that you want to build so the the the main St machine is quite uh simple in there and then we have some um conditional jumps uh somehow I mean this this logic there at the end we are just so the operations is a plus b equals c maybe with aari or with a some flag that you where you can do a jump to one line of cod or a different line of cod so this is very a little bit the the the the deepness the big idea of the FK itself okay and as I told you this goes through this through this bus okay and one of the advantages and I want to see here I won't give you details about the why uh this architecture is very optimal for example when you have an um a multiplication so the six processor is sending that and uh the arithmetic is also somehow proving that specific operation in there okay so you can have this arithmetic operation but one thing one optimization that we can do is okay for example imagine there are a lot of uh operations that are uh usual when I'm doing for example 1+ 1 or uh 3 + 1 I mean I'm doing the small numbers I'm going to do a lot of them in the in the processor so instead of proving all of them in the processor I can have um another steam machine that's kind of a cash where I have precomputed operations that are used quite often and then so so the processor it's always it's it's sending the same thing but then these uh basic operations can be sended can be proven back to the bus so that means that for these small operations or for these usual operations the cost I mean the number of cells that we are doing are uh very low okay can extend this idea so we can have this processor imagine that I have a operations I can have a usual operations table I can have operations that are doing so a specific stain machine that's doing 32 bits operations and another stain machine that are doing 64 uh stain machines operations just to save here just to save uh space okay this idea for works for example for memory for the memory memory alignment I mean what I'm doing in general most of the process are doing uh reads and bries that are aligned well no problem the the processor is doing some memory operation and the memory that just works in in 64 bits 64 bits uh words then you're doing a read aligned read you get it from the memory aligned read you get it from the memory but what happen when you are reading an an align it well we can have a special St machine that's solving these un aligning things and actually is okay so you want a non-aligned operation this is solve so the the arrow to the bus means adding to the bus the the other arrow that goes in the other way subtract to the BS so okay I'm proving this but in order to prove this I need to do two R operations and of course some Logic for check I mean for putting these two two reads in the same in the same one and then of course the same memory machine is doing these uh two reads in there this allows us to do this kind of techniques it allows us to do for example continuations in a very easy way way so if I have a processor I mean I have a I can I can I I can divide the execution in different different slots in different different phes okay so it this the first 100 clocks second 100 clocks and so on so I I can I can't divide them but at the end of the processor so I need to connect I mean the the next slot needs to have the same state that when I finish well what I can do is just send this uh state to the bus uh and recover this state from the bus and then connecting all them together okay the final one the first one this is because it's cyclic I mean z in general we wor cyclic and then disconnects the same with uh continuations in the ram continuations in the ram of course is not time frame in the ram you want more address uh face but so I can divide the proof of the ram that can be big in smaller pieces and but the idea is the same I have an State transition function that goes line by line but uh from the one Chun to the next Chun I I can um send this the state of the run to the bus recover the state of the run to the bus the last the last address that you are reading or writing and then you can continue so we can split also for example the ram in there so this allows us to divide all the work I mean there is no big state there is no State machine that you need a big thing so you can divide always the work uh in everything okay here I as I mentioned um this there is a direct conversion this is important direct conversion that means that you can use Ras or or any other Pro or any other compiling language is 64 bits you can use goang or you can use C++ or you can use uh uh even Java I mean there is a so um um uh this 5 64 bits is is quite a standard uh in there and as I told you there is one um direct conversion one to one from uh risk uh uh from risk to thisk we also Al have a kind of a side project that we are um evaluating converting from wasam to CIS and even directly from LM to CIS this is particularly possible and it can be quite optimal because when you go to intermediate code U you don't need these registers and then you can uh have better code in the processor but in general um the cost of the proof um of course the processor is important but I would say that the operations that are doing the processor are even more important in general so also the the mean the well it's not that you are optimizing the processor and then you get a better proof you need it's the whole thing that you need to optimize okay um well this Advantage we already talk most of them okay it's low latency low latency low latency and uh yeah we integrate with the recursor we didn't talk about the recursor but the recursor is the idea is that you can take the the any proof and then you can uh build kind of automatically uh uh a proof that uh Aggregates both proof for example if you are proving two blocks uh of ethereum you can get a a proof that actually Aggregates this proof okay and we have this recursor that's a tooling that allows you to do that automatically uh uh on this okay so latency how we uh push hard on the latency actually there is four pieces I have not much time but I'm going to try to go fast the first one is the one that you cannot paralyze okay it's like okay if I want to prove a computation I need to compute that thing okay so I cannot paralyze that computation very much okay uh and also if we are emulating for example 5 I do an emulation so here I can go as slow but we are trying this to go as fast as we can currently we are at 80 mahz so we can run a processor running let's say the minimum Trace that you need then to generate the winess computation uh we can go I think that there is margin here for at least one order of magnitude Improvement then we have the winess computation the windows computation this is already done in parallel so this can be paralyzed Um this can be paralyzed um uh very much and you can even start building this Windows computation while while you are executing that also the windows computation the idea is that we're doing that work in the worker we are not have like a central Pro we can do that in the worker that means that here we limit a lot the the the bandwidth so we have one bandwidth we don't have bandwidth requirements it's minimal the quantity of information that goes from the uh uh between the the workers or a central I mean the coordinator and the and the workers okay then we have the Su proof building actually this is the high work but the cool thing is that because we can we can we we have control of what's the sides of the sub proof then uh it's just a matter of putting more subs of course if they are cheaper and faster you will need less sus if are slower you will need more and that's very much a matter of a cost but not a matter of a latency because we can control very much this uh proving this proving time okay and then the other last is okay then we need to aggregate okay this is probably cently where we have the bottleneck but uh I mean the good news is that we are getting much much much better and uh this number 10 seconds we already get lower to that here as mentioned plony 2 has very good circuit for there and optimizing this ation circuit is uh super is super important okay there is a talk think I think in 1 hour or so I think it's in this room that will give more details on this U uh uh optimization parallelization all the proofs that I mean all the proof we are been running in a super computer and doing experiments and see that all these numbers get very very good and this scales uh very well okay future plans lot of things to do a lot of things in the bck a lot of optimizations that uh we want to implement and uh you can try it it's working right now we have a basic thing it's working the last time that I said that in the in the in in a in a scenario somebody created a tornado cash and was inter a nightmare but but but just this is work in progress okay so I would say at this point is the status that it's working requires some refactors of some pizzas some cleanups some documentations and we hope that let's see if in the coming weeks everything looks nicer and uh we want to build the it has no PR compiles I mean know these estate machines are limited but you can build any program and you can prove any program for these uh pre compil so for these specialized uh State machines we we we wait uh I mean take maybe uh two three months still to to to to complete okay and that's very much my presentation enjoy it thank you very much let's go through some questions so what issues did you run into using risk V 64 that risk V 32 wouldn't have I mean well these are we have been checking what's better what's worst here there are uh I mean there are pros and cons I would say that the big the big Pros is that when you are doing especially when you are running rust you are doing a lot of M copies I mean you are moving a lot of uh memory information and when you are using 64 bits this is much much much more optimal I would say this is one of the things that that we saw that there is a lot of uh copies uh in uh moving back and forth uh uh in a place and because uh also the the the corrent kurrent uh I mean current systems are 64 bits so everything is much more standard for example if you want to build I mean if you want to use rust is 32 bits is okay but if you want to do go for example you cannot do it in 32 bits so you can do in 64 so 64 is also much more standard and but again the the different test that we did is that we're slightly better doing in 64 in 64 bits so that's why we stick it at 64 beats and how will this be different than sp1 risk zero uh well uh uh we need we still need to do some benchmarks and compare but uh here the thing is that we are designing not for cost I mean that we are here we are going to push uh as hard as we as hard as we can in uh in doing the speeds uh as fast as as as as fast as uh we can so it's difficult it's difficult to compare it's difficult to compare yet also we have some um this uh machines that these special State machines that needs to be finaly but uh I mean the numbers that we have seen is that we are much faster especially in CPU uh we are working very much in the GPU and the the numbers are quite promising but again I don't like to talk yet about the numbers because it probably would not be fair and and and uh things needs to be more um seted uh I would probably both sides to start getting some real comparations from a developer experience point of view what does the developer need to do to incode the chain rules into zis just write a Ras code I mean just just write a program that do do the whatever rules that's in there uh just a side note here is take care on the security of that is um um when you are proving when you're getting some input and getting some output the input uh you can you can still have soundness problems if you are not writing the right way so you need to understand a little bit how the things work even if you're writing a programming rust because there are some security issues that can happen there but yeah at the end is but at the end is writing a rasco and are there any drawbacks that come with cisk I mentioned this security things that uh but this is common in anym okay you need to you still need to understand a little bit what behind but uh it's much more practical and it's I mean it's it's much easier to to to write a Ras code and out it on a Ras code that not an assembly written in some specific special processor there don't all these improvements by build built specific Hardware are they better than making optimizations on VMS I think that's what that says maybe uh if they do uh we will we will we will we will uh Implement them but they are not hard to build this Hardware I mean this is not like a uh real Hardware that you need to build like uh machine machinery and this is just a program and electronics you need to have them the right way and once you have it you're they already there I don't know what these mean so maybe you could read them out if they make any sense to you um yeah this is um I we would probably need the the context but this is at the end is uh is the position of the this is the position of the registers because we are emulating registers in RIS 5 so we are setting a specific positions in the memory that actually are the registers and this is just defining the this is just defining the constants that defines the the registers in in there there question in there yeah that's very much the the thing I mean I can't we can give you the tail I mean there is on that but that's yeah that's the developer needs to write uh P do not at all not at all unless you want to build maybe some specific State uh some specific uh State machines I mean some co-processors or things like that that then you can extend that way but the idea is that the the final process the final person that's writing in Ras just write in Ras and that's it great give a hand uh for Jordy BINA
