# The immediate next steps of ZKML - Cathie So | Ethereum Foundation

- Channel: [ETH Belgrade Community](https://streameth.org/eth-belgrade-community)
- Date: 2023-10-07
- Duration: 32:11
- Watch: https://streameth.org/watch/yt-my_mGaIeWdE
- YouTube: https://www.youtube.com/watch?v=my_mGaIeWdE

## Transcript

um so thank you very much for coming to this talk I know ckml is kind of a Hot Topic right now uh so hopefully everyone is Keen on knowing more I was trying to keep this presentation quite high level I want to get a lot of use cases get all of you in the audience to be excited about ckml and at the end of the presentation then I would kind of talk about like if you're a developer if you want to get involved uh what should you do or kind of like what's the next steps that everyone should be thinking about ckml so this is the rundown as I've said just now so the idea is um hopefully everyone already understands what zkp is or what's their knowledge proof can do and the uh when we see ckml we kind of assume that we already have a CK circuit that can perform some kind of machine learning operations so in this case in my diagram is a CK circuit that can perform neural network so NN stands for neural network right so a lot of the models that we know stable diffusion large language models they are all neural networks so let's say we already have a CK circuit that can perform that then naturally there are three types of data that go in and out of this circuit right so for inputs on the left we have some input data for stable diffusion the input data would be the prompt right an example prompt is something like oh I want to generate a image of an astronaut riding a horse right so that's my input data for a face recognition model the input data would be the image of my face and then the other type of data that goes into our circuit is model weights right so nowadays if you open source models model weights is public for open AI for instance the model ways are private we kind of know how that model works but we don't really have enough data to know exactly how it works and then the output right so again the output depends on the use case if it's a facial recognition model the output could be an encoding of my face right like a hash or ID um for some kind of for stable diffusion models the output is an image right so it all depends but that's how in general a ckml design would look like so from here there are actually three types of use case we can talk about so first use case essentially the use case where we want to keep our input data private but our model weights public right so the way to think about it is there is this model that is completely public open source that everyone can trust right so a model that has very good performance that everyone can trust so that when we put in some private data and we get a result people can trust that result as being reliable right so put to put it in a more concrete use case let's look at this example so at authentication right so a use case with private private input data and public model would be Authentication an authentication that uses your biometric data right so it could be say a smart contract wallet with biometric authentication right then your model would be either something that can encode your fingerprint very nicely right securely Collision resistant if we have such a machine learning model which actually we do right we have used um uh actually a lot of different machine learning models for fingerprint recognition for a long time well then if we can kind of write that model in a Serial knowledge circuit then we can use that for smart contract wallet Authentication well of course there is a lot more kind of complications and kind of imply the things that you need to take care right such as you don't want people to be using the same scan twice so you probably have to check for that but that's kind of the essence of it all right another thing is probably you can use your face and any other kind of biometric data the other use case I could think of for kind of the case for private input data and public model is proof of humanity right so if any one of you have done the I think get passport uh um proof of humanity you know that you have to submit an image and that image is not private because when the machine determines that you're not a human that image of you actually gets sent to some human committee Court to determine whether this person is actually human or not but let's say if we have a machine learning model that can identify whether this image is a real person versus I don't know an AI generated person well then maybe there is a way that we can have a proof of humanity without actually sending our picture off to some server or to some service provider right so that's kind of like the use case for the first type of use cases for ckml the second type of use cases is the flip side of the first type right so we have the option to keep our input signals private or public so now let's switch what if we have private model weights but public input data and the use case of this is you can think of it as more commercially actually a lot of web 2 companies probably are not willing to share the model right so for example open AI has publicly said that they're not open sourcing their models because of a lot of different ethical concerns security concerns so this use case could actually be helpful to brain web 2 models or web to use case onto the blockchain right people tend to think that the blockchain has to be all transparent but in this case if we can at least keep the model weights private which is sorry model weights private which is the kind of the most expensive part of a machine learning training process then there's also a tons of use case for that so for example a way you can think of it is we can have a ZK version of kaggle so for those of you who don't know what kaggle is kaggle is actually a bounty platform such that companies can post their heart problems onto those platforms and requests for submissions from the public to to to basically solve that problem right so it could be classification problem then the metric will be accuracy it could be a generative problem then accuracy is something else but the idea is there's a bunch of people in the public competing to submit their model their best model for a particular problem the problem with this is that actually in order to allow this platform to evaluate whether your model is really that good you have to submit your source code right so of course you can say that well there's legal license like your source code is kept safe unless you really like win the thing but essentially we're still giving up our source code to this platform so imagine in a case where we have ckml then there could be a way that we can share a model well not share a model to a way to prove our model performance without sharing the model right so for example I have a data set of 1000 samples I would just run through this 1000 samples with my model construct the CK proof and submit the proof right and by submitting the proof I'm showing you that my model is 99 accurate okay and that's probably good enough for a competition right and only after I win this competition then I will decrypt the model for you or actually share the model weights with you so this is a way where there's could be um you know private model transaction like how can I sell a private model to someone reliably how can the other person trust that this model is actually you know as performance as they say so this is um one way to do so and uh you know to get it with some design of smart contracts you can actually also guarantee like whether this whole transaction is fair to both sides of the buyer and the seller and the third type of use case is uh actually um kind of counter-intuitive so a lot of people think about CK as being privacy related uh and this actually I added a slide probably within the last month or so to my presentation um because I think there's actually a lot of potential to just make everything public um so you might think that well why would I want to make things public uh you use CK for privacy right especially ckml I probably want to keep my personal data private or I want as a company I want to keep my model private so why would I want everything to be public well we wanted to for a compression purpose right so it's a sinus um if we keep everything public there's still a use case or usefulness to it because we know that currently at least on evm it's impossible to put the full machine learning model inside a smart contract right it's limited by contract size it's also limited by the operations that you can actually perform inside a smart contract but if we have this in ckml then essentially what we can do is we can take all the computation or the machine learning operation off chain right we do everything off chain we have a huge server they you know crunch the numbers computers EK proof and only with the very small size proof we submit it on chain and that could be useful for uh quite some purpose so the one of the case you can think of is uh you know alcohol Trading right let's say if you have an investment down that wants to uh use some kind of algorithm to make investment decisions or actions right how can I prove that this is actually based on my algorithm other than someone just randomly submitting actions onto the chain this could be one use case the other use case or the other application for this particular type of use case is a new kind of idea which is aigc nft so aigc nft is not a new thing a lot of our nfts nowadays are actually already algorithmically or AI generated but there's a way there's always a talk about well how do I prove that this image is actually generated by this person or by this model right let's say I put in a prompt into stable diffusion how do I prove that I actually did that action not that I just downloaded a random image of the internet and minted it as an nft and this is actually a newly submitted uh EIP if anyone wants to do go check it out so the number is 7007 the idea is that well what if we can then prove the process of this AI generated content right so aigc um you know so let's say I have an image that I generated from the proms um an astronaut riding a horse right this whole Pro and then I get an image right nowadays if you want to do an igc NFC you will probably just mean that uh image as an nft but ckml essentially can prove this whole process right so I can prove this process that I actually turn this prompt into this image via this particular model and this actually allows us to do two things first of all is verifiable creatorship right so it proves that I am the person who created this image right so it will solve quite a lot of like IP um I guess arguments of whether you know the person actually was the creator of this image and the other thing is it actually creates a new concept of what it means for aigc collection so in this case one particular version of a model right so let's say I don't know stable diffusion 1.4 uh with a particular seat number that particular model would be a nft collection right because with a fixed model one prompt can only give out a particular image right so that gives you a quick kind of deterministic way of determining you know aigc nft ownership and creatorship so if anyone wants to check it out as on the is still in the pull requests of the EIP Repository and this is kind of the whole process I'm not going to really go through it too deeply but the idea is you can think of it as if you are a user who wants to make use of this particular model that some service provider provide essentially what you would do is I will propose a prompt all right I would say that I want to generate an image with uh you know 100 people in Serbia right in front of mp uh the MPS theater right and then this prompt will go into the model and also there will be approver that generates the proof and then so now you don't just need the image to mince this nft you also need the prompt along with the proof and the image or in other any other form of media that your AI model and your machine learning model generates to right and this generation process is verified by a CK verifier contract and then you mint it as an nft so this is like the third class of use case and it's also the most importantly available use case as if we are willing to make everything public there's actually a lot kind of um uh uh I guess the the proof size that we're talking about is actually quite small and um well basically you get a lot more benefit if you're willing to keep everything public and just want to use CK proof as a compression or accessing this right so uh now we're going to dive in a little bit into history before we talk about the Outlook so just want to kind of present to you that we're still at a very early stage of ckml and for those of you who would like to get involved there are two main problems that researchers or developers are dealing with most of the time so the first thing is seeking proof is in fixed Point arithmetic or modular arithmetic and typically machine learning is in floating Point all right so uh we are actually gaining advantage of like gradient boosts and stuff from the floating Point weights so there's this quantization uh problem that we need to solve right how do we turn something that is in floating point in machine learning into fixed point or modular into ckp and the other problem we encounter is also that we are not getting a a lot of size or depth in our machine learning models currently that we can fit into ZK so just a quick kind of one through actually the term ckml started two years ago so not very long ago and the first example was actually just linear regression so linear regression means that you just have you know a sample of data you determine the best separation of your data right and then split them into class 1 and class 0. uh and then we started thinking well why can't we put neural network into CK maybe we can do that right so then people started putting it first maybe just the final layer and then eventually the full uh convolution Network so that's my Repository and then we kind of run very quickly in the last few months or so so first of all uh in if San Francisco uh the same group that coined the term ckml also did a proof of concept of aicc as nfts and then we also get a lot of performance boosts as well so there are two very important ongoing projects that you would absolutely love to would like to check out if you're interested in this field so see conduits eckl and also Daniel Kang's ckml repository so those two are making use of you know the Russell language and the Halo 2 prover to gain a lot of performance boost and fit like much larger and much deeper model into a CK proof so the latest updates at least as far as I remember is that Daniel's Group has already been able to get GPT too uh into ckp of course where when we say we got gpt2 into CK proof we're talking about using a very very powerful machine with a lot of memory right so it's not like you can do it on your laptop but at least there's already attempts that you can do it and it's on some uh I guess consumer grade available Cloud machine right so well the reason why I want to talk about history is because I want to tell you then what's next right what okay let's say we're already given all the history just now we already have some ways that we know how to make big models available in ckp well what are the immediate next steps so the first thing is actually there are the projects that I've mentioned is pretty much it so that's all the projects in the ckml field so it's exciting but at the same time it's a little bit sad because we do need much much many more projects into this field so if you uh you don't have to be a ckp engineer to get started actually if you have some machine learning background some AI background that's even better and there are quite some open source libraries that you can contribute to I've listed two of my own so that's a machine learning library in circum and also a library that transpiled from like tensorflow Clara's model into circum the ckp language and the idea is why do we need contribution the way I think about it is that we have I think about the team that builds tensorflow or Pi torch right think about how many people they have think about how many types of operations they actually support essentially for all these libraries we have to transform every single of those operations into some kind of a template in CK right so that is not uh you know the the size of the work that needs to be done it's not something that can be covered by just the few teams and me personally so we actually need a lot of people who are excited who knows machine learning or even just linear algebra to contribute to like just moving or migrating all the machine learning operations into CK uh the other thing is actually um use cases right not just build use cases but brainstorm use cases right so one of the biggest problem we have right now is that what we all know this is a very important thing right that's what you have been hearing about ckml everywhere but we don't know what it is good for right or at least we know a lot of moonshot applications or moonshot ideas but we don't know what we can build right now right and with uh this aigc nft idea that is probably the most one of the most tangible use case right now and I'm definitely personally looking for someone to build the first kind of use case I mean I could build a model myself but I don't really have a use case to to um to I guess to issue my own nft collection at this point right so actually some of the projects told me that I actually want to issue a collection for my users and maybe AI generated Arts could be a good thing and you know if you want to build something like that feel free to reach out and I think this could be a very cool kind of like immediate thing that people know that oh ckml is actually useful and the other thing is there's also another class of tangible use case right now which is gaming in particular there's a game that is built by modulus lab called Leela versus the world you can just search it and that's probably the first result you're gonna get from Google and this game is interesting because it kind of simplifies how we can use ckml and build a very interesting and fun use case so the idea is they have this offline Bots that can play chess very well right and this Ai and there's also the other side of the player is the user so the world so everyone can submit sorry submit a action or a step to play against the spot right and the reason we know that it's about not another person on the other side behind the screen playing against you it's because of ckml right so I actually can reliably know that I am playing against something that is controlled by a particular model right and that's what makes it interesting because you know that you are actually playing against some kind of quite good performing model and you're trying to bid the model right so human versus AI essentially so that's also a very tangible use case it's already on chain online and you know this doesn't limit to chess we can do it in with a lot of different games it's just a very fun way so that people can notice ckml so with that I'm gonna close up my presentation and you know welcome any questions I really hope that like there are more people who can get involved in ckml you know building applications you don't need to be a CK circuit engineer there's always a lot of tools toolings and libraries to help you to build that use case thank you [Applause] so I'm guessing it's question time now is there any question I can't really see there's a question there so you mentioned how when there's a private model and public data sort of thing right you can prove that something was generated by a model but are there any ways to sort of fake that you know to make something seem like it was generated by a model to pass a proof but not really be generated by the model um so according if it's a I guess a sound a ckp it shouldn't really happen so I think it's also a good time thanks for the question it's also a good time to clarify a bit so at least right now for right now this circuit in the middle the structure of itself is not private okay so actually you can see the model architecture the only thing you can hide is the weight okay so that's already a lot of kind of uh constraint or limitation to it so you at least you know that this is the structure of the model and all the computation actually the the computation that is being done is public of course in the future we're also talking about like what if we use a commitment of the function or of the circuit instead of the circuit itself right so that we can also hide the model But to answer back to your question so if there's a sound ZK proof it shouldn't be able to forge it of course currently there are still a lot of different little bugs in a lot of different proving system but like once those are fixed like if it's a sound and complete ckp it shouldn't be able to forge like that proof does that answer the question yeah yeah it does and these are non-interactive proofs in that case yeah yeah so all the proofs that we I've mentioned here are non-interactive so we are just submitting like a single proof with a on-chain verifier to be able to say that this proof is true versus this proof is false that's that's really wild yeah thanks do we have any more questions yes yeah thank you uh I have follow-up question is there any way to establish maybe quality of this model without like revealing all model rates for instance if I train the model using like entire validation set and like without training data so I like focused only on validation set right my model it perfectly work but it doesn't work anywhere else so like yeah Beyond this validation right right so the way you can do it that's actually a very good question as well so it's the same use case we're talking the same type of use case we're talking about here right the question is well if you give them a testing or validation data set they can essentially train their model on this set and perform very well right so the way you can actually avoid that is after someone say that okay I'm done training the model you ask them instead the model weight is still private but you ask them to hash the model weight right so it's the good old commit review scheme right so you ask them to Hash the middleweights and make that a commitment either or the smart contract or just publicly broadcast it and then you send them a testing data set right so then whatever proof that they make the new proof it still has to come out as the same model weight hash so you know that they haven't changed the model weights just to get better performance on this new data set that you give them does that answer your question so we have time for one more question are you raising a hand oh microphone please that's okay um the question sounds so silly I'm kind of afraid asking it um couldn't I build all these uh these ideas by just having the model sign the result instead of going doing all the dance with these your knowledge proof so what what does the proof give me the simple signature over the result doesn't so I'm not sure if I understand fully so you by signing uh you probably just saying that like this is a result that's given by me right like so for example I I okay yeah I think I understand the question so let's say um well so the thing is if you can't prove that uh it comes from a machine learning model essentially if you give me a data set I can manually like let's say if it's image I can manually just look at it and be like okay this is a doc this is a cat this is a rabbit right and I can actually submit that uh that result online sorry on chain and get 100 accuracy but that doesn't make the result useful of course if you just want someone to do that particular task right then that's perfect but if you want to have somehow have this model to maybe your company has some other use case that you want to use this model for right then you probably need this model and probably need to before you pay for it get approved that it is a useful machine learning model that I can use it in the future I guess if that answer your question I think so thanks if you want to clarify your question further on feel free to do so no no I don't want to embarrass myself any further so we are out of time and but if anyone wants to any of course if cat is open for another question um I think we can pull it off so if there's any more questions feel free to raise a hand and if not yes uh hi just a short one in the case of these models how big are the execution Trace matrices just the order of magnitude if you have it um I'm guessing by execution Trace you're talking about like prover time or cover computation the intensity so for gpt2 uh I believe I I am not a team that does it but from the team that does it they're saying that you need a machine with in the order of like tens of terabytes of RAM in order to perform a the whole like proof trace of this gpt2 model so that's the the order of magnitude we're talking about at least for now okay thank you very much so we're gonna like I first of all want to go to thank Carrie for an amazing talk and I want to thank everyone for engaging with your questions so can we please get an Applause for caddy [Applause]
