The verifiability vision by Jens Groth | Devcon SEA
Devcon·Tue, Oct 7, 2025, 12:00 AM
Speaker
Imagine all data was guaranteed to be correct. We could build a trustworthy digital world based only on correct data. In this presentation, we will sketch layers and techniques that can realize this dream, in particular proof carrying data and succinct proofs. We will also discuss the connection to the proof singularity vision for Ethereum as well as highlight caveats that apply; humanity is still in the early stages of the journey and there are obstacles and constraints to tackle Speaker(s): Jens Groth Skill level: Intermediate Track: Applied Cryptography Keywords: Scalability, Vision, ZKP, proof, succinct Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/
Transcript
[Music] uh yeah so I'm going to talk about uh the verifiability vision uh so so the idea is that we want to verify uh the full Digital World um and I think it's something that resonates uh with with many of you right this is also what we hope for in the blockchain world that we can uh create sort like very trustworthy uh data that people can rely on so the very this slide here is sort like the very short summary of what is the verifiability uh Vision uh so uh Alexander KZ and a and troma in 2010 introduced the notion of proof carrying data so the idea is that you have uh data and data comes with a proof of correctness and whenever you compute a new piece of data you sort like take verified data check their proofs of correctness and then you know that you're relying on on verifiable data you can now compute your new piece of data and you can create a proof that this corre computation is correct right and that proof would both show that you computed correctly but it will also show that well all the previous data the inputs you had were also provably correct right so you have this recursive composition proes that Mees that the whole graph of data that you're Computing is verifiable okay so what are examples of things that we would want to Claim about data right so so imagine you're you're reading your online news site right you may see an image and it would be nice to know that this is actually a correct image not some some fake uh information um and nowadays there are some cameras that sign the data sign so like what are the GPS coordinates what is the time and so forth right so as an witness to this being a correct image you may imagine using the raw data from this camera and and the GPS coordinates and the signature and so forth right and then the proof might demonstrate that you know this 100 kilobyte image that you're looking at is actually derived from a much bigger set of raw data with these kind of uh meta information associated with it other examples if you building online uh voting systems you may encrypt votes for privacy right but then you know these system can get messed up if somebody submits a invalid vote so there you may have proofs that the encrypted vote is valid um and and here it's actually important to have confidentiality right that you don't want to leak anybody's vote you just want to make sure that they casting a valid vote um so like the poster child uh example that we use for Zer knowledge proofs is that identification purpose right where customer wants to prove they're over 18 years old and people are now s like actively going out and getting data from say passport ships or something like that so you have some raw material to work with to give these kind of proofs and in the blockchain world there's a lot of applications right so for instance you may want to prove something about a smart contract execution okay so let me sort of like generalize that a little right so cryptographers they then think about this as a class of statements okay and talk about proof system for these class of statements and there we have typically a setup that we expect to be publicly available to all parties um sometimes there's some trust assumption in that so there may for instance be a ceremony in which you could create that um but if that is once that is done then it's there and available to everybody and then you specify a pro algorithm a verify algor the pro algor would take the pro key the statement you want to prove and potentially a witness if there's such a thing to the statement being true and produce a proof and the verifier conversely would take the verification ation key the statement and the proof right and then either accept or reject and this was invented back in in the 80s where um the definition was that should be complete so if if it's a true statement then the pro has a witness should be possible to create a proof and conversely if it's a false statement then it should be infeasible to create a proof and then the twist in the story back then was that well it turns out the proofs can actually convey that well this is a true statement without conveying any other information right which is very surprising so the Z knowledge property is that you know well the proof reveals nothing about the witness it just gives you one bit of information namely that the statement is true so it's this kind of proofs that we would want to give for the data that we want to verify what I want to talk about today are what I see as as challenges for proof carrying dat to really take off um and so like four aspects of that I want to talk about the visi ility and ease of use of these proofs the cost that comes with with them uh to create them um composability right because the idea of proof carrying data is that you you actually compose proofs you would have like inputs and they have some proofs and you would want then to compute new data and give proofs for those and so you have to have some compos posibility of of the proofs and then I'll talk a bit about the trust anchors how do we link these proofs to to real world data okay and so so this is like a bit of a a history of like you know how did people go about as sort like EAS of views for for developers in the olden days and this is an old picture of me I uh did my PhD work working on voting systems where you needed to create proofs of of valid and back in the day the way you construct a proof system is that you hire a cryptographer and then they construct a special purpose proof system for you um um and then in the recent years there's been a lot of uh exploration of building tools that make it easier to construct proof uh so something we could dub as s like the ASC approach is um that you have a program you want to prove that you know inputs to that program yield a certain output well let's compile the program into a proof system that can handle that and more recently uh there's a lot of teams now working on Z kvms where this is sort of like giving you a more traditional experience of proving something so the idea is that you take a program you compile it down to something that the VM can execute and then the VM takes this executable file the program p some inputs and produce an output and together with that a proof that the computation was correct so let's dig a little more into the so like what does that look like um so um so ZK VM takes as input a program um at Nexus we're using risk five as the target programming language but of there also many many teams building ckvm so you could also find this for for wasm or special purposes language like Kyu and and so forth then the program can take an input uh an input X they may also take an A Witness W and will produce an output Y and then appr proove Pi that this output is correct and there will be a matching verify algorithm that then looks at the program and and the inputs X and the output Y and the proof and decides whether it accepts this proof as being correct and the whole goal for for all these teams that are building ckvm is to give a very easy developer experience right so so the way it will look like for a developer something like this you can just use a command line and you can prove or verify these computations so I think we're getting somewhere with sort like the the ease of use from a developer perspective um the next thing the next barrier I think is is the cost so let's sort like dig into what where we are right now and I think the whole Community has settled that that the type of proof that you want is a snark okay and what is a snark well it's an the one of these proof systems where we want the proof to be succinct okay so the succinctness property is that the proof is really small okay and we're talking about in the best cases something like a couple of hundred byes okay and the interesting thing is that it can be a couple of hundred bytes regardless of how complex the statement is right so imagine you have program it does like lots and lots of operation but you still just get a very small proof and something that's very easy to uh to verify uh we want the proof to be non-interactive um and to give a little bit of context because that's not present on the slide here so in the olden days when people were thinking about Ser knowledge proof they were thinking about proofs where you had an interaction where the verifier was interrogating the the prover um but the nice thing if it can be non-interactive where the Prov just produces the proof and then it's done um is that now you can take these proofs and you can copy them and you can distribute them right so if you see a proof uh for for some claim then you can just make a copy and you can send it to somebody else and they will be convinced as well and then we want to it to be an argument of knowledge which is sort like a technical way of saying that um the only way you can construct a proof is if you were also able to construct a witness so now other words you're ensure that it is a true statement okay so when we have succinct proofs right then we have these very small proofs and the verification time is really small as well it's independent of the program execution time it's also independent the verifi does not take this auxiliary witness as input right and that's super cool right because it means verification can be much cheaper than doing the computation itself and I think this is like a a key message here right that verification is cheap okay and that's what sort like makes the verifiability vision possible Right In some sense all these data that come around with proofs right the proofs are for free right so it's like from a user perspective okay I can get a piece of data and blindly trust it or I can add exactly the same cost almost get the data and the proof that it's correct right which one do I prefer well if it's for free I would want it to be verifiable okay so where are we today uh proof size is really small it can be as low as a couple of hundred bytes the verification time is really low as well it can be as small as just essentially reading the statement what is the bottleneck nowadays is the proving time and and that's horrendous right so the overhead of like you know the cost to prove something whereas you just compute it directly are it's like many orders of magnitude there's a lot of work going on to improve that right uh and if you saw like look across time right like back in the 80s when when Z proofs were invented people didn't even think about like the efficiency of the pro it could have infinite time to prove something and people have become more and more careful and I think now we s like you know the way you measure it is you would talk about like you know how many cycles does it take to prove a cycle uh so trying to compare apples to to to Apples um and in the future I mean there's a lot of companies working on Hardware acceleration so forth right at some point we may sort like go to using physics terminology to to measure something like you know many how much energy do you spend per per approved cycle for instance so we are at sort like an interesting phase right now where we have these proofs that while they're efficient to verify and they're efficient to copy and transmit they're really expensive to create right and the question is who's willing to pay for that okay and and it's very nice I think zero knowledge proofs have been fortunate to meet the blockchain space because there are actually a lot of applications where we're willing to pay for that in the blockchain space right because and that could be because it unlocks a unique capability it could also just be because there's so much replication going on when you're verifying computation right so think about thousands of validat nodes that have to redo the exact same computation to verify that it's okay on a blockchain right well instead of doing that why don't we just send a proof wrong that's very easy to verify that hey this is a correct next block and also uh the cost of storage is something that can be prohibitive on on the blockchain right so instead of storing a lot of data and and so forth now maybe we can just have a proof that says oh there is some data that would result in this end state so one of those examples right ZK rollups right where it used to be that you do unchain execution so you have this kind of like low throughput expensive operation whenever you're processing transactions and now we're talking about well why don't we do all of that off chain right and we just create compute what is the the Delta what is the change that would happen to the smart contract from all of these transactions and then create a proof that that's the case and then we just give the the change and the proof to the smart contract and can update the balance okay um so the next thing I want to talk about is composability so how do these things glue together and here it's really the succinctness property why we like snarks right because they ensure that these recursive proofs of proofs of proofs and proofs that you have when you do proof carrying data and you're Computing new data based on verifiable inputs that they stay small right because all of these proofs are small and whenever regardless of how many inputs and so forth you have when you're Computing new piece of data you can still compute a small succinct proof for that being a correct operation so the nice thing about snarks is that they they compose well right they're succinct so there's no size explosion in the proofs they're non- interactive so you can sort of like give them as inputs and be used in further computation and the soundness of them is well cryptographically speaking uh close to 100% so also there's no degradation it's not like you are getting sort of 9% and then you have like another 10% loss in the next step and so forth right you are staying close to 100% soundness And this is unlike other types of of evidence you could give right so for instance if you want to do direct recomputation to verify something well then you would have to send all the data right so there would be no succinctness anymore right if you wanted to do some interactive techniques or like interactive proves or spot checking or something like that well there's some interaction you can't do recursion you can't give that as input to the next step and for Economic Security uh techniques which I think is a really cool idea right but it becomes difficult to manage this kind of Economic Security and track that over a long path of of statements one thing that's interesting about um about this proof cost then when we use it in proof carrying data is that we can amortize the cost over many verifications right so if you have a proof that you've created that suep and then you're building on that and you're building on that and building on that somehow you're acre value of verification value in all these subsequent steps right are relying on that proof and I think this is sort of like really sort like nice thing right but it also points to so like maybe there's a threshold effect that the more data we put into a proof carrying data system right the more we can amortize the cost of of creating a proof over many steps okay so so the last part I want to talk about um are um trust Anor and Trust management and in blockchain we have the Oracle problem right how can we get data into a blockchain and the same thing happens with these proof carrying data right sometimes there are data that com some from from outside and we just have to trust them okay and how do we incorporate those and that's what I refer to as as trust anchor right so there's a question of how do we incorporate that data so maybe it's a camera that signs a raw image it's e theum that finalizes block and you have some sort like threshold signature multi signature on it or something like that right and this is an adoption barrier so so there are not many places in the world today where we get these trust anchors for free people have to go around hunt for it but the exciting thing people are going around and hunting for it and trying to strip you know data from from passports and using that for identification protocols and so forth it does lead to point to a problem as well though right which is proof management what happens if a trust anchor fails okay now you create a lot of data on top of that you have proofs right and maybe you lost sort like what was the origin of that so what if it turns out that well somebody who signed something they actually malicious and and they signed some fake input so I think there are lots of these kind of management issues that we still have to resolve with s knowledge proofs and proof carrying data and I think that's sort like going to be something people are not thinking too much about now but I think in in the years to come that's going to be an issue that is uh increasing in importance okay so I want to to wrap up so like talked about some of the barriers I see and so like where we we're going and I'm hopeful that we're we're going to get there so so what is what is the end result what does success look like um I think in the blockchain space uh we're talking about the proof Singularity uh so think about taking all the computation and take it that offchain and doing it where it's cheap outside the chain and really using sort like the L1 chain as a verification settlement all right because that's what the blockchain is good at it's good at verifying stuff and more broadly I hope we will s like be able to go so like beyond the blockchain space or maybe we can hope for the blockchain space becoming so like all income presenting is something that's very widespread in use and there are what I'm I'm the vision I'm hoping for is s like a digital world where we have this web verifiable data right and since verification is cheap right hopefully people would want most of the data to come with proofs of of correctness right and we would have a much more trustworthy World um and with that uh thank you for for your attention all right thank you Jens uh we're actually running a short short on time but we can take some questions uh type theories have thought a lot about composition of proofs how can type theory help with composability of cryptographic proofs okay yeah um huh um I'm not entirely sure okay so but but I think I I guess in I suppose it's been from a somewhat different angle that type theorists have been thinking about this I guess in terms of like uh so one thing that you're using type theory for is so like static analysis like so will a program execute correctly will it process the data correctly um and that is an orthogonal thing to what we are trying to do with with these kind of proofs which is more like has it executed correctly so so in some sense I think that is sort like a complementary thing and I think one should do both right so you should use sort like formula verification and type Theory to make sure that whatever you're doing is correct and maybe for proof carrying data there's also a question of like how how does one computation Carry On from the next one and does that make sense and then you would want the proofs to make sure that well once you know that that is all like the correct structure of of data and computation then that that has actually been executed correctly uh what are the advantages of elliptic curve based proofs versus small field proofs yeah so I think elliptic curve proofs are um I think in some ways easier to reason about you can have some very nice homomorphic commitment schemes like Peterson commitments um the the disadvantage is that you're working with larger fields and that can be a costlier operation um so if I had to conjecture I think the world is moving toward using small Fields rather than ectric curves just for performance reasons and then potentially also for postquantum security all right and a lot of people want to see this answered should we be prioritizing postquantum proof systems uh yeah um so I think it's not as urgent for proof system as it is for say encryption schemes because the proof systems if there's a proof there suin they don't actually carry the data around right so it's more a question of like how long time do you want the verification to hold and nobody could break break the soundness but for anything you verify here today right when since as far as we know no quantum computers exist that should be good enough um I do think that the world is probably moving to post towards post Quantum security anyway just because of like they seem to be pretty efficient techniques anyway for creating proofs uh what proving system does Nexus use yeah so we have uh several approved systems so we have um a ckvm uh We've uh implement reimplemented the Nova family of folding schemes from the ground up so that's Nova and Supernova and Hyper noova we have in the ckvm um but we also have an implementation of of jol in in the ckvm all right given the time constraint I think we can only take one more uh what's an application of this verifiable layer you would be excited for that is encour possible yeah um so so I'm very concerned about deep fakes uh and in general so like the the news that you get and the fake information you get misinformation uh which is very deliberate misinformation that we're facing so I would love to have um a system by which you had much stronger guarantee whenever you were reading an article or piece of of news that you know that there sort like a path back to the original sources that provided that information um I've s like tried myself talking to journalists being misquoted not maliciously it's even worse if it's malicious so if there's some sort of way of saying you know in an article guaranteed well that the original person who they talk to agree that this is correct interpretation so forth I think we could all already eliminate a lot of misinformation and I think this Zer proofs can be part of that uh I also think that there other components uh involved in in getting to that solution all right so unfortunately our time is up but thank you so much for the amazing talk J [Music]
Automatic transcript — names and jargon may be misspelled.