Advancing OP Stack to ZK Rollup: Achieving Efficiency and Security with Zero Knowledge Proofs
Devcon·Tue, Oct 7, 2025, 12:00 AM
OP-stack based rollups now retrieve L1-to-L2 deposit transactions and L2 transactions from Blobs. Current solutions face two issues: 1) increased operational costs due to batch submission overhead or 2) protocol complexity during challenges. We'll share our experience addressing these using ZK fault proof. Our new challenge system is cost-free for users and easily extendable to ZK rollups. The presentation includes our example of switching from zkEVM to zkVM and optimizing proof generation speed
Transcript
[Music] [Music] uh show too many slides um hi how we guys doing so I'm ta AKA fake death I'm a ZKA engineer from chroma so today I would like to share our experience of advancing all P tag into zare rollup so we're trying to enhance efficiency and security with zero knowledge proofs so for those of you who don't know about chroma I want to First celebrate our successes from last year our Mainland launch so uh uh we were one of the projects built on op stack so we're one of the first I mean we're the very first optimism rollup optimistic rollup with an active fault proof system so we utilize CK proofs for it and so ZK enables to solve some problems in traditional fault proof systems and first we approach with a circuit circuit based implementation of ZK provs but it has a very very big issue so we're trying to Transit from circuit based approach to cdk VM based approach and lastly we all we're also building a ZK backend library to push the boundaries of proof generation speed so why ZK fault proofs So In traditional uh fault proof systems it requires various of it requires a lot of interactions which leads to high bond amounts which decreases the level of decentralization so with ZK fault proofs we can just narrow it down to a block level instead of instruction level which requires much more much less uh interactions so uh it not only reduces the operational cost itself since there's less interaction so we need less computation but also it can significantly reduce the bond requirement so it it naturally leads to better decentralization and security of the network so more over for like arbitrum spold if we apply ZK ZK technology into it we can we would be able to sign reduce the bond requirement significantly that is usually used for uh blocking Cil attacks also optimism is uh putting a lot of resources into exploring ZK fault proofs so we have a uh announcement that we have received uh retro pjf round five from optimism to explore more on the ZK fallproof size and this recognition I think uh shows how much important and uh possibilities are there into building a permissionless fa proc system with ZK so we've looked at like our achievements since last maintenance launch and now I would like to share some our lessons learns and challenges that we've experienced for building uh CK fallproof systems in our main net so the very important fact that we learned is that the circuit based approach is just just not sustainable so we currently have about like 100k L of code which are for custom circuits so we have to write it in a circuit language like plish and it's very very hard to uh write and debug and verify so what we're trying to do is check the Integrity of evm St State transition function but we have we need like several uh specific custom circuits written in circuit languages to verify that the state transition function given the input is done correctly and back in 2022 uh vital given mentioned at the time there was no like uh production ready zkm circuit codes but they had a POC version of it in ethereum's psse and back in the days it was like 35k lines of code but even vital mentioned that it's never going to bug free for a long long time so it's very prone to errors very prone to human errors so to sum up there are two challenges writing the circuit itself is tough because it's hard to audit and debug and also since optimism and ethereum keeps keep brings upgrades to the main net to support like full compatibility we have to keep up with the protocol upgrades they're making which means that we have to modify circuits again and again every time the upgrades come up and this could lead to like possibly if we make more modifies modifications then there will be possibility of more errors which would like harm our security of our maintenance so we need a better approach so now we're trying to take a ZK VM based approach because ZK VMS can give us generality and auditability what it means is as I mentioned in a circuit based approach we need to write circuits in a circuit language like plish which is very unfamiliar and because of that the Auditors will have also have a hard time auditing and even us can't really verify well that our code is written well but ZK VMS allow us like flexibility we can write any program that we want to verify the computation in Rust and since it's written in Rust it's very very better to audit and debug so now I would like to dive in more to ZK VM area so how we're trying to move on from circuits to ZK VMS so as I mentioned ZK VMS can provide a general purpose environment for verifiable computation so uh what it means is that we don't need to implement circuits anymore we can just write rust program and ckvm can basically execute the program and give a proof which proves that the computation was done correctly so how it works is first the each zkv M vendor would have a compiler tool chain that I can compile a rust program into let's say if it's based on risk it would turn into a machine code based on risk 5 and then inside of ZK VM there's an Executor which executes the program and generates what's called an execution Trace execution Trace is basically a polinomial Bound by some constraints so back then we have we would have to write like circuits to represent some input data into a polinomial and then have constraints for it but now zkm circuits could handle that so now we can just write rust code there's more flexibility and then in the proving stage where the ZK VM gives a proof based on the execution trace it would commit to the trace and then generate a proof so I have a like a summary kind of uh picture that we can compare like how it's different so the colored green parts are the code that we used to maintain and we have to maintain so back then we would have to maintain like zkv ZK evm circuits written in plish so there would be like evm circuit RP circuit MPT circuit and transaction circuit which could uh which kind of replicates the behavior of evm so it will verify the logic if the let's say the transaction body and like transtion I mean like the block header and block body comparing with like hash it and then like compare with the block cash uh that's what we used to do it's very very tricky but now since the on the right side there are ZK VM circuits so ZK VM circuits would handle that and instead we would just write a blog execution program which would which for us as a ethereum layer to we would derive the batch from L1 and then execute it on L2 so that will be the program we were trying to write and uh some of you might ask like vitalic mentioned in 2022 we launched our main it in 2023 and then like why didn't you guys first approach with ZK VMS at first right like why did you guys ride circuits and do all that hard work that's because back back then uh ZK VMS were not like performant so since the proving time is long in in within the challenge period we have there can be possible volum possible attack Vector for delay attacks so that's one of the reasons why we chose a circuit based approach but now uh RIS risk zero calls it continuation and like sp1 calls it sharding like there was a breakthrough called sharding which with the full execution Trace you can now divide it into a a unit called shards so it's usually represented in the number of cycles of risk five so sp1 they kind of have an optimal number of like two to the^ of 22 cycles for each shart and since we can divide it into charts now we can uh parall generate a proof for each Shard and then compress it later Into A Single Shard so now there's more more uh upside in terms of throughput since we can paralyze proofs so that's what we're trying to do and we're almost there shipping it to the main net and now now I want to like share more of our of my thought that what we can explore more in the future to overcome like barriers to become a from a optimistic rollup to a zare roll so my concerns about zkm based approach is since zkm vendor uh vendor business model is usually running a prover Network so they get a delegated proof and then they generate a proof for it and then return the proof to the client but there might be a situation where zkv Andover Network might fail so uh usually zkm vendors since they've built their ZK VM they're very very good at uh like giving a good performance so usually they would have like a cluster of machines and then they would like uh split the workloads to each machine and then like orchestrate it and then create a proof and return but since most of the such Prov Network related codes are close sourced uh maybe to handle this kind of network reliability issue we might have to we might like us chroma might have to implement uh a code to orchestrate multiple machines so that every validators could run by run I'm not sure if this is correct but they they might have to run a cluster to generate a proof for the network so uh another concern might be so now then now there's going to be more stage one rule ups but for stage two uh we we we only want the uh Security Council to come in and resolve like the problem if the failure is only onchain provable but currently in my perspective for the ZK VMS currently there's no way to make like the network failure to be onchain provable so that's that's enough another concern and also since we cannot Overlook that the fact that ZK VMS itself might have a bug uh so what I think is we need a multi-pro system so that we can compensate the bugs for that so then in that case the goal would be to enhance the security of the network but not much with the overhead so we don't want like uh periods to be longer so there might be like three types of multi-prover schemes instead of like having a single prover so one I think most of the networks like scroll and Tao is trying to use teth but in my opinion uh there's a trust issue that we have to trust the te implementation itself so and there's also like Hydro centralization problem so we might have to come up with a clever idea to maybe use like three vendors of te's and then like have a multiset kind of thing and another one would be kind of very very labor intensive but come up with another circuit B another circuit based ckvm so instead of plish and Halo 2 you might have to write another ZK evm code circuit code with another language uh I don't think no one will try that and I think the best would be uh our approach to have another ZK VM based zkm I'll tell you why like why we think this is the best uh and if if we have a multi-prover system and if we have like ZK VMS now we could really really easily ship ZK rups then like pass one or two years but I've did some benchmarks with several ZK VMS but there's still a feasibility issue so first there's a cost barrier which is very very important and in my opinion it would be really really hard to like charge the users to use it since like most of the layer twos and most of the optimistic rollups now have very very cheap uh fees so there are three types of costs that we need to pay to run a zq rollup first is it's the biggest bottleneck which is the proving cost so uh I've tested with like a network with about like three TPS and you need about a million dollar for proving the blocks for a year so I think it's not that feasible now and also we would have to com commit to the batches we've uh settled into layer one so there will be settlement fees also and to finalize the blocks we would have to send the proof to the contract and verify on chain so there's also going to be verification piece but that's going to be really really small because we will be able to agre aggregate proofs inside a ckvm which would be way cheaper but there's going to be trade-off between finality and also we could Explore More on future directions where I introduced the multipro system and also uh the ckvm vendors would have to consider more about Pro decentralization where they don't keep the Clusters running on only their Network and getting proof requests but instead of like have a incentive mechanism so that Pro Network can be decentralized a bit more uh and like we're also pushing the at the same time pushing the boundaries of proof generation so we have our own very unique uh open source Library called takon which is a particle that's known to be f faster than the light uh and it's it's basically a modular ckvm CK backend library and we also have a feature of Z GPU acceleration so since we've built our ZK VM based on halo2 we first optimized Halo 2 a lot and we have about like 1.5x 6X faster Benchmark and we also are trying to utilize sp1 ZK VM at the moment so we also optimize planky 3 as well which is used by sp1 and last since circum is very very widely used in the community we also uh optimize a lot in terms of Rapid snark so it's also faster and we could wrap up into this one figure where if if chroma has a multi ckv Andover back uh backing the network we would also have takon which is a modular ZK backend Library which could support all all kinds of ZK VMS so we would only have to maintain the very front like ZK VM main which is written in Rust and which mainly uses the library Kona which is main from maintained from optimism which has block derivation logic and execution logic so we could easily write uh a program for a main net within like 200 lines of code and we would would have an input with which the which would be the block we're trying to prove and we would have a multi-prover like one of the Frontiers risk zero sp1 o VM uh Alida maybe and then we have a takon which is backing with powerful GPU optimizations which will accelate accelerate the proof generation speed yep that's it for today and you can contact me with fakee 9999 on Twitter Telegram please contact we could share more about uh fault prooof systems I would like to hear your ideas and that's it thank [Music] you oh you can just stay over here all right uh great talk to so thank you to the crowd for submitting all these amazing questions uh we do have less than five minutes to go through some of them uh let's go through uh the most upvoted one how is it different than op6 sync yeah uh some of you guys might know and not know but uh opung and us we have a very very similar logic actually we were working under water like for about like two months and then like we shared it to the sp1 team about oh we have a choma SPO improver repository you could take a look if the logic would be correct for ZK VM based Prov and they were also working on water with the same thing and like it comes out that like we have almost similar logic and they have an additional feature of what I mentioned of aggreg aggregating the proofs so the logic is very similar but basically we built the same thing uh so I'm going to mark this as answered any stats on Pro cost specs and time so the Prov cost for uh the for our fault proof system for about a block with a TPS of three to prove a faulty block uh of tps3 I think it costs about like uh currently if you so it depends on the pricing of the N the approver network you're trying to use but if you're just using an ond demand plan which would be the most expensive per Cycles uh if I remember correctly it's it was about like 20 20 bucks for a block with like about like uh tps3 and uh the their the stats and like the machine specs it's not revealed at the moment for SPO Andover Network so I'm not sure about that all right do you know if the ARB ecosystem have similar plans yes I I definitely know and I also uh know the fact that they have a fast withdrawal feature now with ZK proofs I didn't have much time to like deep dive into it but I I definitely think like arbitrum also is going through the right path so maybe we could collab more and talk more about it all right great so we we can go through I guess two more questions real quick uh in your prior implementation by pish circuit you mean Pony 2 or three or something custom oh it's it's actually uh Halo to's plish arithmetization uh it might have like confused you a little bit but it our we're using a back end currently Halo 2 all right so one last question what happens if the guest program panics can you still produce a proof so if a guest program panics then uh you can't produce a proof so uh I'm not sure uh the person who ask who's asking the question like is intentionally panicking the program or not but like yeah if if it's a bug that's causing panic then proof is not produced so this is one also another thing I mentioned in the slides before is we need a way to like uh like overcome like if the program panics like but if it's a bug inside a ckvm or if it's a bug in a program then like you would need to wait to uh verify it on chain or like have a way to prove it on chain so that so that Security Council could like intervene or what all right great so our time is up thank you so much ta uh
Automatic transcript — names and jargon may be misspelled.