New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

A Deep Dive into ZK Proofs of PODs | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

Provable Object Data (POD) is a format any app to easily sign data and make ZK proofs without manual circuit writing or trusted setup. Proofs are described in a simple configuration language, then compiled to a family of General Purpose Circuits (GPCs). POD tech is used in Zupass and available as open-source libraries. We’ll dive into the cryptography behind PODs, and what makes them suited to ZK proofs. We’ll also cover the modular design which makes GPC circuits highly configurable. Speaker(s): Ahmad, Andrew Twyman Skill level: Intermediate Track: Developer Experience Keywords: Libraries, Zero-Knowledge, Cryptography, zupass Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

[Music] all right thank you again everyone uh thanks for everyone who stuck around for the Deep dive session um I know these are being recorded So anyone who's watching this session online you may want to go back to that one first it'll give a a more gental introduction um this session should be a little bit less rushed because we have a bit more time but you know it's definitely going to get a little more Technical and I hope can answer some of the questions that you may have had after the first session um so we're going to cover two things um I'm going to do a deep dive into how pods work how we build them into a Merkel tree what makes them provable um and then how that kind of proof comp proof compilation step uh happens that turns things into a circuit uh and then my colleague Ahmad over here is going to take over and talk to you a little bit about the actual circuits underneath that how we make them modular how we make them uh make the these very flexible proofs okay so let's review very briefly so we're building this e pod ecosystem um where issuers issue credentials or attestations um that may have come from some non-z friendly data source outside of our uh ecosystem um such as like the pre-ex data base that Devcon tickets come from um you're going to hold them uh in zass or somewhere else uh you can use pods in whatever app you want zass is optional um but we let users hold their own data um and then that that user can in response to a Quest EST make a uh proof about that data in order to prove you know I'm a licensed driver I'm over 21 I have a ticket to Devcon Etc um so this is what we're trying to enable um so pod is a data format that makes ZK proofs easy um we use merization or rather we build a Merkel tree I don't know if merization is actually a word but I will continue to use it um we do it in a deterministic and repeatable way and we use this data structure so that individual entries are easy to make proofs about without having to prove every entry of the Pod at once and I'm going to in this session go a lot deeper on how that works um the other thing that makes this easily provable is we use ZK friendly Primitives um all of the underlying math is based on uh the baby Jubjub Prime field that's those like prime numbers that were uh or numbers modulo a prime number that I referred to earlier um we use a Poseidon hash and eddsa signing um these are two algorithms that can operate over this uh baby Jubjub Prim field um and the proof compiler plays an in important role which is that there are some things that are not ZK friendly to verify but as long as they are public and be verified can be verified by the verifier directly um they don't have to be in the circuit um and I'll show you a little bit more about how we make that work um so a reminder of what a pod looks like I've now I've cut this pod down to only three entries just so that everything fits on the slide but it's otherwise the same as the driver's license I showed you before it's got a name it's got a date of birth um and it's it's got the public key of the card holder um so how do I sign this in a ZK friendly way like right the normal way you would sign a Json you have to if you have to just sign it like a string um and you have to be very careful that it's all in the same order otherwise your signatures is wrong um so we've got a better way in pods that makes it much more provable so what I'm going to start with is I'm going to take all those entries and the first thing I'm going to do is I'm going to alphabetize them um to make sure that they're always in the same order um entries actually come from a pretty or entry names come from a pretty limited character set they basically match an identifier in most programming languages it's like asky Plus numbers and underscores um so there isn't any Unicode question about like what the alphabetical order is it's pretty clear um so we do that we put the names in the values adjacent and then we hash them all um and in this diagram um and in the future diagrams green is something that is ZK friendly I can do it inside of a circuit if I want to Red is something that is not ZK friendly I can't do it inside of a circuit at least not efficiently not in a browser on a phone um there are certainly ways of doing non-z friendly things in circuits they're just really big circuits so all the name hashes and any of the string values that I want to Hash I'm using uh Shaw 256 that's not a ZK friendly hash function so that means that those strings are going to be represented in a circuit only by their hash not by the actual string because the circuit can't repeat the the hashing um I mean that wouldn't be very CK friendly anyway because strings are variable variable size and ZK circuits don't like variable size things um but this is the approach that we taken um the green P's here that's the Poseidon hash so uh a date is a number so I can use the Poseidon hash function to make a hash that is ZK friendly um the same is true of the public key a public key is actually two members so the the P here is actually a two input Poseidon hash rather than a one input Poseidon hash the diagram is lying a little bit but it's still ZK friendly it's something that I could verif verify in a circuit later if I want to okay that's step one I haven't built a tree yet I've just done a bunch of hashing so let's start building a tree um we're going to Hash together the adjacent elements um that means that the name and value that correspond to each other always hash together to get a new value these ABCs at the top those are just the hash that results in hashing together the previous two hashes um okay so we're starting to look a little bit like a tree let's keep going next level up I'm just going to keep hashing together adjacent elements this is a pretty standard Merkel tree um there is an open question here of like what do I do when my tree isn't a multiple of or power of two um the way we do that in our the particular version of the mer Merkle tree we use is we just skip that level so when we go up to the next level we just take that value C we feed it up to the next level unchanged um this is a data structure called a lean IMT IMT standing for incremental Merkle tree um it was built for semaphore uh by the psse team and we're just reusing it here um it is a very lightweight Merkel tree and works particularly well for small trees um most pods uh right now or less than 16 elements or you know less than 32 elements you can go as big as you want but the circuits just got a little bit bigger with the depth of the tree so the lean IMT help works very well and very efficiently for small trees um okay so now that we've gotten to the top we've done all these hashes we've got one hash and that's called the Merkel root um in uh Merkel trees that's a thing that you might publish to prove that you're a member of a group for instance um in the case of PODS we call that Merkel rout a Content ID that's the the thing that I mentioned earlier is going to represent this pod the same content put through this same process will always generate the same content ID so we we make this a first class concept that like you know what the pod's content is this is its hash essentially now that we got the content ID um that's when we can actually sign it and as I mentioned earlier the content ID is actually the thing that's signed um you're not feeding in the the names and values directly to any signature algorithm you're just doing this merization and then feeding the content ID into the signal algorithm um you're also going to feed in your private key if you're the issuer and you're going to get out of signature um EDSA signing you notice that that's kind of a gradient between green and red um it is a ZK friendly signature scheme but it turns out that there's a key generation phase in the signing step that isn't ZK friendly so from your private key it's actually hard to generate a signature in a ZK circuit that's not usually a problem for us because all you need to do is verify it like if you ever wanted to do the ZK Circuit of equiv valent of signing you would just sign outside and then stick the signature in and verify it in the circuit and you're good anyway but I'm trying to be correct about my color coding here because I'm That Kind of pedantic Okay so we've got this thing that represents the entire pod um how do I prove one one entry of that pod that's the thing I most often want to do in one of my ZK circuits so this is what Merkel trees are for if you've seen them before um this is called a a technique called a Merkel membership proof that if I want to prove that I have this entry called date of birth um all I I don't have to give you the whole tree I just have to give you uh the path up the tree and give you the sibling that corresponds to the rest of the tree at that level and if you're a verifier and I give you all of the values that are shown here but not the ones that were deleted um you can verify that all this is correct right you can you can take the the name and value and hash them you get B you can then take a and hash it with b and get X you can take Y which gave you and hash it with X to get the content ID and you can validate the signature so here I've just like avoided having to give you all of the elements but you can still validate that this is a valid pod so that's pretty cool um let me formalize this so this is the the standard Fields you'd have in what's called a Merkel membership proof for an entry proof um so the root is the content ID the leaf we always uh set the the name of the entry to be the leaf that's a little bit arbitrary um the depth in this case is three in order to validate the proof you have to to know how deep the tree is um we've also got what's called an index this is telling you when you get those siblings along the way are they coming from the left or coming from the right um that's because our hash function when it takes two inputs the two inputs are not reversible uh so in order you have to know which direction the tree goes to go up so that's what the index is for um you have to get those S sibling hashes um and this is an interesting thing that I didn't realize until I built this the value hash that corresponds to this name is right there in the list of siblings so it doesn't actually have to be fed into the circuit as a separate uh element it's just part of the meracle proof it has to be there anyway um which will be useful later uh I've got the value and then I've got my public key and my signature for verifying um so this is a this is what's called a Merkel proof um it's not a ZK proof um when I built this I initially thought oh well if all you have to do is prove one element of a pod one entry I should say because that's the word we use um I could just send you the Merkel proof and I don't have to make a ZK proof at all wouldn't that be easier it's certainly a lot cheaper um anyone here have an idea of why I can't do that why it would be a bad idea why it might be unsafe uh well I mean disclosing the signature doesn't necessarily obviously mean that uh I'm disclosing anything like what would the attack be can you can you answer that oh yeah yeah the siblings yes so what would you do with that if you were an attacker yeah precisely so these are hashes it's hard to reverse a hash but if the values themselves come from a relatively small set you can use a hash to try and guess what the value is right so this is much like a rainbow table attack on a uh on a password list like if you leak the hash of something it is much easier to find out what that thing is than if you don't add that hash so that's why we don't usually send these Merkel treat uh proofs around directly outside of a of a circuit thank you for that um and that was something I actually learned while building this I'm like oh that would be easier wouldn't it well it's easier because it's non-secure as it turns out unfortunately um so the next step here is we put this inside of a snark so now instead of drawing a diagram of like the work I do in my code this is now a diagram of an actual ZK circuit um but it's very much like the the previous one um and then on the right side I've basically grade out all of the fields that I'm not revealing one of the magic things about a ZK uh proof is that I can choose which inputs are public and which ones are private so the only things that have to be public are the name hash of the leaf because you have to know like you aren't going to accept driver equals true if it you don't know that the name is actually driver so that's always public um and then the public key that signed the Pod has to be public so you can verify that um with some exceptions and if you have a more complicated circuit you don't have to reveal that you could instead prove that it's part of a a known list or something like that but at this level it's always going to be public everything else is private um did you notice the trick that I pulled there I feel like you should probably be asking questions right something disappeared um so the name hash disappeared um that's because that's an Noto ZK friendly hash have an question uh we're not using katak we're using sha 256 but it's also non DK friendly um it's because it's a string and a string being variable length is not going to be ZK friendly anyway um at least within the limitations that we set for ourselves like you can prove things about strings there are ZK friendly ways of hashing that um but it's not very friendly because of the variable length so um for Strings we always use the the shahash um and we do it outside of the circuit this is why or another reason why the name of an entry always has to be public that means that the verifier can take the name as a string the verifier can uh hash it using shot6 get the get the name hash feed that into the circuit and be confident that that hash really does correspond to the name driver or uh I guess I guess it was Data of birth in this example um yeah so that's the other thing that like is important here that there is a little bit of work the verifier has to do outside of the circuit they're not just going to check the ZK proof they're going to check some of the inputs to make sure that they correspond usually just by hashing them again they're not really checking so much as recalculating okay so the last step this is now a proof of one element of a pod um but it's not modular I said that GP GPC circuits are meant to be modular well the way I make it modular is I just slice it up um so this in a in one of our GPC circuits would be represented by uh three different uh modules the object module which proves that a Content ID is properly signed that proves a pod um an entry module that proves that this entry uh or this name exists in this pod and then a numeric value module we call it because we can only prove the actual value in the circuit if it's a ZK friendly hatch function which only happens if it's a number so the numeric value module lets you prove that the value exists okay let me review what we got through all of those those diagrams um so an object is made of entries which are named value pairs the name and the value are represented in the circuit by their hash which might be a Poseidon hash or might be a a Shaw hash um and we use the pints that come in with the data to know how to Hash the value is it an integer is it a date which I'm going to convert to an integer is it a string um is it a is it a set of bytes is it a public key Etc um we generate a Content ID by building a Merkel tree um and that means that I we can make proofs about individual parts of it um and we make a signature on the content ID to make an adastation um so the p and pod I usually use it to mean provable it also can mean another thing which is portable um one thing that I like like about the data format that I described and the process I described is that there's nothing in there about like a string representation or how you have to canonicalize it the way you would sign a Json it's all about math so if you can do that same math and put get together the same content ID then you have a pod and it's valid um that means that while we do support a Json format for sending pods around because it's convenient you don't have to use that you can save it in a binary format you could use a a protuff uh the Frog crypto game has a different Json format that's a little more bit more nicely readable for players and it all is just as valid a pod okay but a pod also is provable um it allows for efficient ZK circuits because of the the hatching function and signature function we chose um in the ZK circuit there's a fixed cost per pod to validate the signature there's a fixed cost per entry to validate the entry and then a ZK uh snark can verify all of this um this is the full list of value types that we currently support it will probably extend I'm not going to go through all of this in detail um mainly they are what you would expect um a few interesting details um any value can be compared for equality because we just look at that hash and can compare two value hashes together um some values will be be equal to values of other types for instance like the Boolean with a value one and an integer with the value one are actually the same value so keep that in mind if you're ever using uh this in your in your uh checks um but but only certain values can be done can be compared in an ordered way meaning less than greater than etc those are numerical values right you can't do a less than on a string in a GPC circuit just because the string is only represented by its hash whereas for a numeric value you can prove that oh this is the value and I can check the hash in the circuit at which point I can feed that value into further logic to check greater than or less than um but also because of the some of the limitations of what you can do efficiently in our particular uh circuit language comparable things are limited to 64 bits in our circuit that's an aspect of the circuit we could make a module for comparing things that were bigger we just limit it to 64 bits whereas a cryptographic number can be anything that fits in a field element you'd use that for a hash but you're probably not going to check those for greater than or less than okay um one last bit of detail getting into uh how we compile down to circuits before I hand off to Ahad to explain those circuits um this is the diagram I showed before I'm not going to go through it again um but I want to talk a little bit more about the compile and decompile steps um that happen in this diagram where we Bridge into the world of ZK circuits um so your GPC is a modular circuit it's made up of modules that are basically sliced up the the circuit that I showed you earlier U plus other things like right I showed you the object the entry and the numeric value amod is going to show you a lot more um the modules are connected to each other by passing around verified values such as a a value hash such as a name um or such as a numeric value like those are all examples in the previous diagrams of a value that will be fed from one module into another um and there are public inputs to the Circuit that say okay something like uh uh entry module number three is going to take its content ID input from object module number one and that's how you would how you would wire these things together that say that oh the first object that I proved has this entry and and I want to make those correspond to each other and then it's going to validate in the circuit that that content ID really is the same um so if you think in terms of Hardware there are a bunch of like uh multiplexers and Dem multiplexers that are feeding things around uh if you think in in terms of uh software it's a bunch of like if statements essentially um and all of those wiring signals or configuration signals are by by definition public um they won't be hidden because the verifier has to send in the same configuration to make sure that they're verifying the same proof okay uh and lastly as I mentioned the GPC uh circuits come from a family we we Val we generate circuits with different sizes for instance a typical circuit that we use for for Ticket proofs has one object a Max of five entries two numeric values on which you can do things like less than or greater than one list that you can check membership in ETC um and we generate a whole family of these things with different sizes you can just pick the one that is cheapest for what you need um and you have to download the artifacts for the circuit that you picked and this is like a vague size range um proving artifacts are pretty big at least on the the level of a a proving on a mobile phone um verification artifacts are very small so if all you have to do is verification it might even be reasonable to just build all the AR effects directly into your your app bundle but for proving you probably don't want to do that you probably want to put them somewhere else okay um this is kind of a restatement of this so the proof compiler takes the the configuration takes the inputs um including the pods the user SEC private key um and list things like lists and Tes that you can generate um uses them to compile down to the proof outputs using uh there we go uh the proof compiler so what this is going to do um much like uh a a programing language compiler it comes in multiple phases first thing we're going to do is check all the inputs that they are valid uh in by the definition um this includes some things like range checks that are actually important for instance if you're going to prove that a value is in a range between Min and Max our circuits only work if Min and Max are 64-bit numbers and not bigger than that um because of certain assumptions that so that means both that those numbers are always going to be public because they're part of the configuration but that the prover and verifier both are going to check that to make sure so that they can trust the output of the circuit um while checking the configuration the compiler is going to determine what are the requirements right how many modules do I need in order to generate this proof um it's going to pick the smallest circuit that fits uh alternately you can optionally just feed in a circuit identifier if you want to say always pick this circuit please because that's the one that my verifier knows how to verify um then it's going to hash and meriz all the inputs following the procedure that I showed you earlier um in order to format them and the configuration as input signals to the Circuit potentially a lot of them if you've got like a list of 200 entries and have to see feed them all in um but usually it's in the it's in the tens of inputs not the hundreds um you're going to download the artifacts um the proving key and the witness generator in this case so that you can then finally generate the proof and then you're going to decompile some of the circuit outputs back into uh the uh the revealed claims um and some of this logic actually has to bypass the circuit right if if what I revealed was a string and what comes out of the circuit is that String's hash I'm going to go look back at the inputs to get the actual string because I can't reverse the hash um but the prover has access to that so you going to just copy that string into the revealed claims okay um the verifier is going to go through all of these steps in a very similar way and that's because the verifier is double-checking that the prover did everything right um all the stuff that happens outside of the circuit that is security relevant the verifier has to repeat because you're not tusting the pro right things inside the circuit are being verified by the ZK proof itself but anything outside of that the verifier just repeats the work including things like checking that the bounds are within 64 bits and they can do that and then the proof will be fine um the only difference which I've called out in bold here are that the verifier will rather than picking a circuit it confirms that the circuit that was selected actually fits the inputs um it downloads the verification key instead of the proving key which is much smaller and then it can verify the proof um and then as I mentioned in the in the over intro session there is some work to be done after the verify compiler and the verifier is done right the verifier tells you this is a valid proof with this configuration you should make sure that that's the configuration you asked for you might do it by just before calling the verifier like replace the configuration the prover gave you with the one you expected and that'll just work you copy over the circuit identifier to do this we have some examples um but you might also do something more complicated where you say like this isn't a configuration that I will accept it might actually be stricter than the one I asked for it might be slightly differently formatted that's up to you um and then also you should check whatever's in the revealed claims they were revealed for a reason presumably so you should check that that's something you actually want to accept and go do whatever you're going to do with it in your application okay uh that is it for me my throat is feeling a little bit parched from all this talking um so I'm going to hand off to Ahmad now who's going to take you into how do we actually do what's inside of the circuit and and he's gonna plug in his own laptop to do so e e right all right so we'll be going through gpcs in a bit more detail so when I initially prepared this talk I had some background slides on pods maybe I'll do a very quick recap of what was covered in the last couple of talks but the tldr is that gpc's allow us to make ZK proofs about pods so that's you know essentially that's the point we're trying to drive home and if you want to look up the on GitHub here's a QR code to our repo so we've got a mono repo um pod package GPC compiler package GP circuits it's all in typescript um there is an ongoing uh let's say initiative to do more in Rust um if you actually look for the parket repo ZX Park parket actually should have put the code up um then you'll find uh the beginnings of a rust implementation of uh much of what we've done on the Pod side at least not the GPC side yet so yeah we already know pods are cryptographically signed key value store has anybody read the blog post by Joel Gustafson does that ring a bell merizing the key value store for Fun and Profit I mean yeah this isn't an entirely new idea but it is quite uh quite solid cryptographically um Andrew's already gone into detail on how you know we meriz a pod but you basically start with the Jason object um you appropriately meriz it you sign the Merk root and Bam pod so we've already used terms like content ID so that's the uh root hash um the signature is the EDSA Poseidon signature associated with the Pod um yeah and then we basically have a series of entries so key value pair is the name of an entry the value of an entry so that's the terminology we use um just to go through some examples because I am going to talk I'm going to show you some GPC configs corresponding to certain examples but imagine a world where we had pods as ID cards you know ID cards with pods underlying them so here would be one such example we use Jason because Jason is the accepted uh I guess uh data representation format in this uh in this world um we have I fields in our ID pod like name date of birth semaphor ID uh like you know public semaphore ID um so we basically take something like this the state would actually issue a pod with this underlying data and you know once a PO is constructed you basically have the content consisting of this part plus the signature plus the signer signer's public key this would be the state that signed the ID pod um and the content ID which is the mer route that can be reconstructed in a deterministic way because we already have conventions about how you arrange all of these entries um and then another example would be a ticket to an event such as this one so a ticket should have an event ID so the event you know the tickets associated with a product ID like uh you know regular ticket maybe a speaker's ticket a ticket ID to identify the individual person and then maybe the attendee has a semaphore ID well the you all do via zoo pass um so that would also be embedded in there and in much the same way we obtain an appropriate JavaScript object according to the current implementation that represents our pod now life before pods was quite complicated as far as ZK proofs are concerned you'd always have to Define an appropriate uh data structure choose some signature scheme find a way to just cram that data into a circuit right and you're gonna have to probably pad things along the way do various little ad hoc things necessary for your application glue code and then who knows what comes next and eventually you profit hopefully and that's how it was before pods um and after pods things get simpler so now I'm going to talk more about the ZK side of things how we can like um what what sorts of proofs we can make about pods um with our GPC framework general purpose circuits um so basically we have a many parameter family of reusable circuits so depending on the proof you may or may not need to prove about just one pod multiple pods maybe pods with many entries pods with you know five entries whatever um we need various circuits to accommodate you know all of these possibilities so many parameters um and yeah each circuit has a number of modules with a specific purpose that I'll go into in a moment and this is how we make proofs about pods so here are some examples that relate back to the pods I showed you a second ago so you could say you could you want to prove to somebody that you possess an unredacted speakers ticket to Prague crypto 2023 which was part of Dev connect so as you can see I haven't updated my examples but it's a Concrete example um another proof you might want to make is that a ticket to a Dev connect 2023 event was issued to me um without necessarily revealing what event and proving that you are over the age of 18 without revealing who you are so these are some things you might want to prove um so we can be more explicit about what it is we're saying in each of these examples so for the first example you are saying I possess a ticket pod it is signed by the organizers of the event so for pro crypto 2023 would have been zerox Park psse um it contains the event ID corresponding to Prague crypto 2023 it's product ID corresponds to that of a speaker's ticket and its ticket ID entry value does not lie in the list of redacted tickets right and you could always have a field in the ticket saying is redacted but why would you ever set that to true so in this case there would be some Universal list um list of tickets that are no longer valid um and then with the second example um that a ticket was issued to you you would yeah show you would say that you possess a ticket pod it's signed by an organizer of a connect 2023 event um and it contains an event ID corresponding to one of those events and you actually possess the private data corresponding to the semaphore ID that's embedded in the ticket and finally with proving that you're over the age of 18 you say that you possess an ID pod it's signed by some trusted State uh it's date of birth entry value it lies in the appropriate range to prove that you're over 18 and you actually own the semaphor ID so it really is your ID pod so from the user point of view this should all be streamlined by the way I mean there's going to be some web page you go to and then you click click and then you get some pop up and it's like yes select your pod select it prove bam you've proved what you needed to prove that proof goes to the server together with whatever you chose to reveal I mean what you choose to reveal is usually determined by the application the developer would have bake that in you you can actually go through and make sure that you're not revealing something you don't want to reveal in which case don't make the proof and then whatever it is you were trying to do hopefully should come through so here's a high Lev description of the GPC compiler as a recap uh as mentioned before from the developers perspective depending on your application you would form some kind of proof configuration um specifying what it is these pods what what properties these pods should have the pods that go into the proof so the application the user interacts with will have both a proof configuration and proof inputs the proof inputs will have all the private stuff that needs to be put into the circuit to generate the ZK proof you need uh the proof configuration will have other details like has an entry called date of birth um maybe date of birth should lie in a certain range or this entry of this pod should lie in a certain list its value should lie in a certain list and uh yeah so you basically take these things um something I'm kind of sweeping under the rug I'll get to in a moment is um the family of circuits I said that there's a many parameter family of circuits thing is we've pre pregenerated a lot of artifacts we're using circom and Groth 16 under the hood um so you know you have to do your trusted ceremonies power of tower whatever um so we've done that for a certain well not a proper trusted ceremony but you know we have some artifacts that you can use um but an additional input to many of these procedures is is the family of circuits you're using so there's a default value here that we provide but you could just as well customize it and the process goes as follows you plug in the configuration and the inputs and you know you put it in this GPC proof box um the config and input are all checked I mean the user might be trying to prove they're over the age of 18 but if they're not over the age of 18 according to their ID pod this will be caught here and let's say it isn't caught here you know you have a script Kitty they basically comment out that line well they're not going to be able to generate the ZK proof they're going to go to cryptic error right or even if they did try to even if they commented out more lines further down the line they're not going to come up with a valid proof so here this helps uh helps you not shoot yourself in the foot essentially in producing false proofs helps in debugging as well so everything is checked all the validity of all the statements the types Etc um and then a circuit has to be chosen you know let's say you have 10 pods you want to prove about yeah to pick a circuit that's big enough to accommodate for that so um we essentially build up a data structure called GPC requirements which is just a appropriate Json that lists you know the various things we need and then the circuit whose size is the smallest as far as number of constraints is chosen from a list um so something we haven't quite done maybe we should think about this at some point is uh who's interested in onchain applications yeah I mean so besides I mean you don't really care about verification cost you probably I mean proof size is not really much of a concern it's more about input size you might want to optimize for that that's maybe something something you might have to like do on your own I believe we have there are a few projects that do that sort of thing but that is another thing you might want to optimize for but in our case we just optimize for um circuit size because then that reduces proving time because proving can take a long time if you just choose one of the one of the larger ones um inputs are compile uh config is canonicalized I should say so um sometimes the proof configuration can contain unnecessary restrictions on the inputs for example you could say that a number should lie in the range you know 010 but not in the range 510 right that I mean it's unnecessary to specify it that way so you could simplify that down to saying it you know lies between Z and five five for example or 0 and four depending on whether you're doing inclusive exclusive bounds um but it basically brings it down to a canonical format and then all the inputs are compiled in the sense that once we've chosen a circuit in a circuit you don't have variable siiz Loops you don't have variable size arrays everything has to have a fixed size so there's going to be a lot of padding involved uh in various places and we generally choose the padding so that everything is still logically consistent for example you could say you have a pod has an entry with a value that lies in some list okay list is let's say five entries long but the circuit you chose says all lists have to be have to have 10 entries so what do you do pad with one of the existing entries we choose the first one that way when you prove list membership you're still proving the proper list membership you wanted to prove right even with the padding so we have to do all that so we basically prepare all the inputs for the circuit and then plug that in snar JS get our Groth 16 proof and we're happy besides the proof we get the config back together with a specification of which circuit we chose along the way and finally the revealed claims are included in that output so basically in our proof configuration we might have chosen to reveal our date of birth and not actually check it in the circuit for example for the ID pod example that would be amongst the revealed [Music] claims and verifying is basically taking the output of that last process plugging into a box the config is checked against the claims you can't check everything because there is going to be private data but you can check whatever is available in the revealed claim to make sure everything everything's consistent and then you verify the proof so just a yeah using snark JS in this case and then output true or false so going back to the examples for example I possess an unredacted speakers ticket to Pro crypto 2023 we basically have the proof input uh sorry the proof config I'll start with that so we specify the pods that need to go into the proof we only have one pod ticket pod doesn't really matter what we call it here it's just a label um and then that should have an entry now so yeah so it should have a it should have a sequence of entries specifically event ID product ID ticket ID we choose to reveal both event ID and product ID why because event ID will reveal what event it was Pro crypto and product ID will tell us what kind of ticket it was um the ticket ID itself should not be revealed so we say is revealed false but we do specify it should not be a member of the list that we call redacted tickets now this list itself is going to be specified here in the proof inputs another thing I stated is that the signer's public key should be revealed um the signer of the ticket because we do want to we do want to actually assert that it's a legit ticket right anybody can can issue a pod you just need to make sure this is an authentic pod issued by the the organizers of pro crypto 2023 um yeah then you proof input would basically just be the pods you make sure that obviously ticket pod should coincide with the one you mean in the config um all this is done under the hood by the way in zup pass when you make ZK proof so you basically get that drop- down box you get to choose the the individual pods and then you can make your proof um the membership list should be provided Again by name um list of redacted tickets and you might want to put a watermark in the proof inputs so thing is you could make this proof and then somebody could take your same proof and try to submit it again right if you didn't have a watermark nothing would stop them from asserting the same thing as you so by putting a watermark like a time stamp or something associated with the session you can you know guarantee can't be reused um so yeah that's essentially what you put into that GPC proof box and then out of that you'll get some gross 16 proof which is some sequence of btes not here and you get the bound config which is just the proof config Json or JS object together with a circuit identify field which contains some label that deter that tells you what circuit configuration you chose there is some uh yeah there there is some logic behind the sequence of characters here's what it says I've chosen the protopod GPC circuit with one o one object 10 entries it has a max Merkel depth MD of six what that means is the the pods you put in are not going to be bigger too big that they can't be accommodated by Merkel tree of depth six yeah so two to the five entries um zero NV means zero numeric values um in this particular proof I don't need to say any anything about any numbers per se I'm just checking for list membership and I'm revealing some values I don't need to say that you know this entry is greater than five or this entry is greater than that entry nothing like that so I don't need the numeric value module and I certainly don't need the EI entry in quality module um which is what would allow me to make those inequality comparisons uh this is something Andrew talked about briefly um in order to say things about entries in the sense of inequalities like greater than less than not uh not not equal to that's a different story but if I want to say one entry is greater than another first I need to check they lie in an appropriate range that they are 64-bit signed integers that's the convention we use um but you do have to restrict the range to do this properly in ZK so I don't need any of these modules I do need a list membership module that's what the L is for I need one list because I only have one list here to to check that the ticket isn't redacted um maybe it has 100 entries you know so I need one list of 100 entries I need zero topples I haven't really combined any entries um you you might want to combine entries to check that a tuple of entries lies in a list or doesn't lie in a list um and in this case because this was Pro crypto 2023 before we migrated to sema4 V4 IDs I need one owner module so that's the O of sem 4 V3 and zero sem4 V4 modules and then revealed claims it's so up there you can see I chose to reveal the signers public key event ID and product ID and they're all revealed here amongst the revealed claims the water mark is always revealed red um because that's you know part of the challenge in a sense and the membership lists in question are also revealed so that's that example um to assert that a ticket to a devc connect 2023 event was issued to me um the approach is much the same the only difference here is I actually form a pair uh consisting of the signers public key and the the signers public key of the ticket pod that was issued and the event ID and I basically check that that pair lies in the list of all Dev connect pairs basically pairs of you know Dev connect event IDs together with sers public keys to ensure that the ticket is uh from the yeah is from that list um and then on the other hand I also assert that the attendee semafor ID entry of that pod is my owner ID and in fact it's of type sem4 V3 you could replace that with V4 if we were dealing with V4 IDs so in this case I think everything's self-explanatory except the circuit configuration changes everything is okay up until about here now I need one Tuple of A2 um so this is like the best case scenario by the way um you might not necessarily get something that exactly fits your your requirements it could very well be that we haven't provided any circuit configurations with one object and a miracle depth of size six perhaps we only provided Merle depth of size five in which case you would have gotten different parameters here but this is just for the sake of illustration and finally we have the example proving that you are over the age of 18 um in this case what you want to do is show that the signer's public key lies is a is uh in the list of public keys of trusted States now in this case I've deliberately chosen a smaller list size 10 I don't know if there are 10 states that we can actually trust in this world um but for the sake of example um 10 um yeah so um things are much the same except for dat of birth where here we assert that it isn't revealed so we're not revealing our date of birth but we are asserting that it lies in the range 0 to 18 years ago specified as unix's time in this example um yeah and uh and in this case since we're actually using our semaphor ID we can actually associate a nullifier with this uh with this um particular proof If we so choose in this case you can specify that as part of the owner section of the proof input and that comes through with the revealed claims so just a quick refresher I've already alluded to this but circuits have limitations compared to regular programs most mostly As far as like um static versus dynamic memory you've only got you've got to make sure everything's of a fixed length all of our functions have to be naturally expressible in terms of addition multiplication in a field this motivates much of our con many of our constructions and much of what we do um and this is why we have to provide a number of templates to accommodate uh different proofs about pods so let's get into the modules and there are so many of them uh let's go in order so we've got the object module each object is as I said um is a merized structure so you basic basically take your key value pairs blah blah blah Merkel rout you sign that you get a signature you have a public key associated with that signature you want to check that all that is consistent in the circuit and we have one module that does that it's the object module this will be done for each object provided to the Circuit like I said we pad things so if you choose a big circuit that requires I don't know 10 objects but you've only got four then one of those objects is going to be repeated and you're just going to have extra checks of the same sort so yeah public key signature content ID which is that root hash all check for consistency essentially just check the signature against the message given by the content ID now each entry you plug into the circuit the way we do this is the entries are their own thing it's not that you plug the objects in in their entirety you don't have to you only plug in the entry that you need to prove something about and then you specify Which object they correspond to so there's an index in there we have a bunch of arrays we have the array of like object content ID and other things um and you specify for each entry which uh pod object underlying object it corresponds to and uh what the entry module checks is that the entry actually correspond to that pod that you specified so you provide a Merk proof of inclusion of that entry so the root hash is recomputed checked against the the object hash and we also um like as was mentioned earlier the actual key the name of the entry is passed into as a hash um you know that's also passed in here to it's also checked against it's in the Merkel proof and in the end if everything is good the revealed entry value hash is sped out because we choose to reveal an entry or not reveal it and by by the way if anything is not satisfied the circuit just won't compile that's what's going to happen oh the proof is not going to compile sorry I should say um yeah now besides entries notice that in those earlier slides I said things about the signers public key the signers public key isn't actually an entry in a pod it is part of the Pod data structure but it's not actually an entry we call it a virtual entry for that reason it is passed into the circuit um it's actually yeah passed in together with the object itself so what happens inside our big circuit is that all of these are aggregated into an array both the public key and the content ID there are some things you could prove about the content ID you could prove the content ID is a member of a list that might be a bit weird because you know who's publishing content IDs um but you can also say things like I have a bunch of PODS and they're all distinct so that's something I'll get into in a moment we actually have a module for that the uniqueness module so we basically take this data and we hash it so that it's treated just like any other entry can be used in all the constructions we have in ZK for proper entries so you can check list membership um you can check for equality non-equality not inequality because these things are going to be pretty big they're probably going to be close to 254 bits right there's no guarantee that you know a hash is actually going to be a 64bit signed integer in a field I mean why would you expect that um so that's what the virtual entry module does um and then for virtual or non- virtual entries we might want to check if they're equal to one another you might um yeah want to check if one entry value is equal to another entry value or is not equal to another entry value for example you could say I have two pods and they were signed by different parties this is how you would check them um and in this case we basically plug in the value hashes we plug in an array consisting of all the entry value hashes together with the index of the entry that ought to be equal or non-equal and you plug it into this thing it'll check for equality and actually it should spit out a Boolean that part's missing here um and that and as part of your proof configuration you would have specified whether you wanted to check for equality or non-equality and then you've got the owner module so if you chose to say that a certain entry in a pod was actually an owner ID semaphore V3 or V4 type then you ought to provide your semaphore secret as an as a private input to the circuit and yeah you basically specify which which of the entries of the uh of the inputs um corresponds to the semaphore commitment you can plug in both a V3 and a V4 identity by the way so it is possible toly have like a pod with a semaphor V3 identity one with a V4 identity and you prove that you have the secrets to both identities that would be a way of establishing that you own both of those semaphor IDs I mean one such way I mean surely they're more direct ways but that is just one way um and uh yeah you basically specify which of the entries does correspond to your semaphore ID you plug in your secret you might want to plug in a nullifier as well to invalidate this proof if somebody else want to come and use it later on um and you might or might not choose to reveal that nullifier hash and in this case either you output that nullifier hasher you output minus one which is our convention for the absence of data um and then we have the numeric value module so here um like I said entry values are passed in by hash by default but sometimes you want to prove things about entry values um like they lie in certain ranges um and in this case they have to be sign 64-bit ranges um so the numeric value module is responsible for that you plug in the actual the value corresponding to the entry that you're arguing about in this entry value hash array um you plug in the entry value hash uh and uh a lower bound and an upper bound and those are checked in circuit and assuming you've done that you can also check whether such an entry is less than greater than less than or equal to greater than or equal to you can do whichever using the entry inequality module so in this case you just pass in a pair of Entry values and this box will just check if entry one is less than entry two and that's sufficient for checking any possible inequality why because we're going to get a Boolean out so you could just flip it if you wanted to do the other way around or you could just assume it's false if you wanted to do the other way around with an with an equal sign as well like a non-strict inequality so in the proof config you don't have to think about that like oh I can only do is less than or is not less than no you don't have to think about it just put in less than less than eek greater than greater than eek it's all the same um it'll all be translated appropriately when compiled down to the Circ so the compiler takes care of all these little things that you you probably have thought about you've probably done this before but it's easy to mess up if you have to do this again and again and you know keep these different conventions in mind then we have the tupple module how do we handle tuples easy we just hash them so in this case uh you might so in one of our examples we had a pair of Entry values and we want to show that that pair lies in a list of pairs so what how do we handle that in the circuit we actually hash the list of pairs pairs we hash the thing that we want to check is a member of this list of pairs and then we check that the hash is a member of the list of hashes and that's what the Tuple module is responsible for it basically just hashes everything away to obtain a tuple hash which you can then you can then treat the tupple as an entry in its own right as far as the logic is concerned here then we have this membership module that one's fairly straightforward we just do a linear search through through the list if we actually uh we plug in the entry value hash and the list of hashes to check lists are always hashed anyway um entry by entry so we're just doing Simple arrays in future we might actually have like you know pluggable Merkel trees so that you could actually do a Merle membership uh but for now we just do linear lists because we only care about small sizes anyway and it's fairly cheap um so yeah the list membership module and then we've also got the Pod uniqueness module as I mentioned you might want to say the pods going to the Circuit are all distinct as far as their content is concerned and this one has a simple switch I mean you just plug in all the content IDs and then you get a Boolean zero or one depending on whether all the the content IDs are unique and whether you want to impose that the content IDs are unique is up to the proof configuration and finally the global slw Watermark module this one just has that Watermark value it's plugged into the circuit a constraint is put in and that's it there's nothing nothing else about it it's just to ensure that you don't reuse the same proof um or just have that the proof be associated with that particular session um yeah so life after pod and GPC should be a lot simpler you just write down the data structure then maybe use the GPC or GPC PCD library or even the zi which uses all the stuff under the hood to you know know put your application together something comes next and then profit and we're good so there I mean future directions for all of this would be more features um yeah I mean I already hinted at a few possibilities like having some form of Merkel membership rather than a linear approach to list membership um more pods more gpcs I mean you know it's pretty the the Pod project is pretty young um lots of frog pods going around so I think uh there's so we're getting more adoption by the day um and possibly a different stack for certain other applications um I mean at the moment we're using uh circom with Groth 16 in typescript um we can't do recursive proofs what if you wanted to plug a pod proof into a GPC well you can't really do that it's not natural in this setting people are going in this direction recursive proofs you know maybe use something like Pony 2 uh or whatever else is in fashion to get things done um yeah those are just some possible directions uh so this is pretty much all I wanted to cover uh what I usually do at the end of or towards the end of such a talk is maybe like quickly go through the actual protopod GPC circom file but I think you can all look it up in the repo um maybe it'd be good to turn to questions instead um I was given a QR code damn it's really small let me see if I can make make that a bit bigger or did you already revealed that QR code no you didn't okay yeah let's do that um should I maybe answer some questions because there have been some questions already so show up last two slides we'll jump to that okay uh so a couple links just to to tell you where to go for more um uh this this was a lot in a deep dive session so thank you all for sticking with us um but there is an all day CLS from zerox Park about programable cryptography tomorrow um all of that I think may be of of interest to you but in particular there's a workshop in the afternoon which is what that QR code is uh that is going to be about zup pass and the sdks and things that you can do to build on uh pods so uh hopeful some of you will join us for that um also check out the docs check out the telegram um and then here's the thing you've all been waiting for we'll leave it up for a minute or so before we switch to questions all right everybody got your frog let's get to the get to the questions okay uh when making ZK proofs are they already being done with the idea that they are anti- quantum or are we still F far from having those problems um we as a team are far from thinking about that those problems because as I said in my intro we're limited to kind of tried and true technologies that live live in a browser and that limits our choices to not really being able to kind of pick and choose that much um everything here is based on uh elliptic curve cryptography uh with the baby jobjob Prime field and uh a curve so I don't think that that's guaranteed Quantum resistance um but I don't think anyone's broken it yet but yes we have potentially would have problems with that in future and that's why we'd want to investigate alternative uh Avenues and proving systems like amod alluded to um could I compare and contrast pods with verifiable credentials um yeah that's something that I looked into it they do something very similar um particularly the uh Iden 3 or pado ID system where they use Json LD to represent verifiable credentials and make proofs about them um it's a similar kind of merization they use a different kind of Merkel tree they use a a sparse Merkel tree which is optimized for larger trees with deeper depth and can do things like prove uh absence of an entry that's one of the limitations of our pod Merkel tree that you cannot prove that an entry is absent only that it's present um if if you saw my list of data types for pod one of those data types is null that's actually there specifically so if you want to prove that something doesn't have any other value you can give it a null value and prove that it is present with a null value rather than have proving absence but for that we get Merkel depths that are like five or six or eight in most of our circuits whereas the um ID3 circuits for sparse Merkel trip trees have a default depth of 32 because they need to avoid collisions on the on the Merkel has um so that's a technical difference we're optimizing for like lots of small objects rather than a big complicated tree is kind of the the trade-off we made there but they are compatible in concept there's no reason you couldn't take a verifiable credential and turn it into a pod pod is much more General like any names and values will go okay um what are some of the challenges problems in this space anything we could look into and help uh almad may have more thoughts on that in some of the like crazy circuit stuff he's been doing like I know that lack of ability to do things like recursive proofs is definitely a limitation for us um uh the size of uh proofs that we can make in a browser on a phone is definitely a limitation for us um we my kind of rough rule of thumb has been we have to keep things under about a 100,000 constraints if you've written ccom code each constraint is kind of a mathematical equation um so uh a our typical circuit right now are anywhere from like 8K constraints at the low end to like 30 or 40K and that's going up to like three objects 20 entries Etc um so you can't really do really really big pods unless you're willing to take some extra time and some extra memory uh it's all possible uh what else is the circuit able to see the raw data are only the tree of hashes how range checks and comparisons done um I think ahmod alluded to that in in uh when he was going through the modules but um only numeric values can go into a circuit directly that's why you can't do like ordering comparison on a string because that would require a string value to be in the circuit but with numeric values you can use the numeric value module to prove that this is the plain text value that corresponds to that hash and then you can do things like range checks uh on that um and there's nothing stopping you from making more modules for more things um let's see uh sorry you scrolled away from the the top um question is there a way to integrate the automatic approver inside the smart CR contract um so proving in a smart contract has the disadvantage that all of the inputs I think are inherently public because that's how uh ethereum works so if there's nothing stopping you I think smart contracts are are turning complete um I think the the size of the input would be very large and you'd have uh confidentiality problems but more useful is verification on chain and we already have that working um using the verif the verify compiler is kind of used as a pre-step to take the like friendly objects and compile them down into the raw circuit inputs and you can send those onchain and have the onchain verifier uh verify that um there's a demo app called frog juice that lets you juice your frogs um onchain uh is it possible to tamper with a GPC config to compile a gener generic circuit that can allow proving false statements um no um assuming that you use the the the GPC compiler as is and assuming it doesn't have any major bugs it hasn't been audited yet um you shouldn't be able to prove false statements what you can do is prove meaningless statements right like I have a pod that contains a field called driver I mean I can make any pod like that and if what someone asked you to do was prove that you're over 18 that is a meaningless statement for that purpose so that's a not entirely solved problem like in the easy case where you where the approver or the the the verifier first asks for this config and theover sends it back it's very easy for the verifier just to say like is that the config I asked for yes or no just like do a deep comparison if you want to get fancier than that like oh you can prove a statement that is like stricter than what I asked for for instance like that's a bit more like logical diction um that's not easy but most cases don't really need that uh the example of a module comparing two hashes why take a list and an index instead of uh just a value I'm not sure what that's about I'm assuming that's the entry constraint module most likely um well I mean it could go either way but maybe you had a particular uh particularly strong response to that I I may have just figured out what this is feel free to jump up if you're the person who put in this question um what the index you're you're seeing may actually be part of that routing Logic the multiplexers that I mentioned so it's not that the comparison is comparing something with a list the comparison is comparing two hashes but the input to the comparison one of them is the entry that this module is attached to so it's always a fixed input the other one you can compare it to any of the other entries so the way that we accomplish that in the circuit is there's a list of the hashes of all of the entries and then you use this index to pick which one are you comparing to so this is that's part of the configure ility it's not about the list being a like value in that you're checking the list is the all of the possible inputs and the configurability is the index uh is there any standard supported for pod um I don't know what you mean by standard so like everything's open source right now the code is the standard I know realize that that's not a great answer um I would like to document this as a proper uh at least a a protocol spec essentially um whether we would standardize it I think is up to adoption and and how important that is like I don't I didn't want to slow us down in our initial development by like getting the w3c involved or the ietf or something like that to standardize things because that can take years I'd rather first build something that works iterate it on it with some feedback um but yeah I think that more of the format should be more closely documented uh that's something that we're going to work on um in order for this to work in the governmental settings do issuer verifier Andover need to agree in advance on the pod's structure um so yes a a pod inherently is schema less you can put any names and values in it that you want but uh there does have to be some agreement between the uh prover and verifier of like oh what does a driver's license look like if you want me to make a proof of a driver's license um we have to agree on what the names of the values are um there are two ways that a a the issuer can actually Express that one is just if the only thing you ever sign with this public key is driver's licenses then the station that is this pod like implicitly is ATT testing that it's a driver's license because this is the public key that is only used for driver's licenses um the other thing that is a recommendation but not enforced by the system is a an entry called pod type that is meant to be a unique string for your app to specify like this is what this pot is for which probably also specifies what the entries should be and what their type should be um but there is nothing inside of our libraries that will enforce a schema um other than the proof config itself is implicitly enforcing a schema of the things that are proven about yeah only yeah so anything that is not mentioned in your proof config is not being cryptographically proven you may be able to trust it because this issuer like states that they only sign things that follow this whole schema but cryptographically if the only thing you prove about is like two entries of a pod and that pod actually has 12 entries like a pod without the other other uh 10 entries would be provable by that configuration like anything that you want to be sure is true you have to prove about um unless you're willing to just trust the issuer um but you can do things like prove an entry present if that's all you want without revealing its value and that is a way of like checking that the schema is present uh but it's really up to the specific uh use case did you I have a question it's answered oh okay any other live questions feel free to put up a hand yeah uh can you pass over the the mic thank you um how do you approach uh proving uniqueness in Dynamic things like ownership of a name or address in a chat room interesting I don't think that's a so directly solved by gpcs um other than you know the standard approach of like pick a random number that's big infinite probably won't overlap but the all you can act I me pods are not inherently unique they are like defined by their contents so a given pod set set of contents will always have the same content ID and then if the same signer signs it it will have the same signature um our signature algorithm is deterministic um so there's no inherent uniqueness there but you can put a unique you know user identifier or something like that in a pod um you can prove in a if you put multiple pods in a single proof you can prove that the individual ones are unique or you can prove that you know the um the user ID that's represented in the Pod is not the same as this other one um but I I think that's about as far as you can go with like this system like proving uniqueness across like all of the users in a chat group or something like that is not entirely in scope but may maybe there's something more specific that I can suggest for you thanks any other questions all right well thank you all for joining us um we'll stick around for a couple more minutes if anyone wants to to chat in person but uh otherwise uh I hope to see some of you at the the workshops tomorrow at the

Automatic transcript — names and jargon may be misspelled.