New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

"Optimistic" Outlook: Pioneering On-Chain AI - Cathie So | ORA (ex Hyper Oracle)

ETH Belgrade CommunityMon, Oct 7, 2024, 12:00 AM

Transcript

[Applause] yeah um unlike the two previous presentation I don't have any math formula as far as I remember in my slides uh so this will be more still on ZK but a little bit more on uh the application side so uh with a pun which is optimistic I put that in quotation marks uh you will understand why in a bit um so we're going to talk about onchain Ai and in particular to bring really large models that we're actually using in web 2 today uh on chain so this is the agenda we'll talk about verifiable machine learning or Ai and then a little bit about zkl just for a context and then we'll introduce something that aura has invented called opml standing for optimistic machine learning and then going into today's U main topic which is a paradigm or a scheme that we call of Pi I'll explain that a little bit as well I know that's a lot of abbreviations uh and then we'll talk about how actually like all these um you know schemes or ideas how does that actually bring Ai and crypto together right so at Aura first of all we believe that verifiability is a spectrum so usually when people say verifiable machine learning they're only referring to ZK right uh it's true ZK is definitely definely sort of the endgame of verifiability right uh the most trustless way of having verifiability but we also think that you can have a spectrum with different security G uh guarantee right so we have CK roll up and op roll up uh why can't we have opml and zkl so on one side of the spectrum we have zkl which is zero uh knowledge machine learning and then on the other end of the spectrum we have optimistic machine learning uh which is uh if you % like op Rob it's the exact same thing right so we are trusting the system optimistically or trusting whatever results being submitted optimistically until someone challenges the result and then in between um unlike row up which is either op or ck well nowadays they're also exploring CK fra proof but unlike a row up where you can usually only use one system to do so uh in machine learning there's actually space for a hybrid of both right so this is what we are sending in between which is uh today's topic optimistic privacy preserving AI so a little bit background on ckl so how many of you would say that you know what ckl is okay so uh ZK guess okay and ml okay great so zkl literally means putting a machine learning inference in particular so we're not talking about like actual learning or training a model yet we're just talking about making a prediction with a model so uh putting it in to a either a ZK circuit or running it in ckvm okay so the purpose of it usually when people revert to zkl there's three purposes which is of course first for verifiability we have a whole slide of that before uh but then there is also sort of two promises that I've put uh with an srisk and you'll know very soon why there's an srisk um so usually people say that hey uh ckl you not only gain verifiability but you also have two promises which is input privacy and model privacy right so you can either choose to keep the input that you put into the model private because of zero knowledge proof or you can choose also to put the uh the model weights that goes into the model private um and before we talk about why I put those in Aster uh even with um you know with all these promises we have problems right which is the same problems that we're facing with zero knowledge proofs nowadays which is it's super memory consuming it's time consuming and hence it's only practical for small neural networks right so to put this into perspective we're talking about proving something like a gpd2 or a llama it's either going to take days or it's going to take terabytes of memory right so and and or or both I guess so that's why um you know there's a a need for alternative uh but not only that not only for all these known issues um actually the two promises that I put in Asis is also not entirely true so we'll look into and this is something that when I was writing the optimistic privacy preserving uh AI paper I was doing some simulations and that's what I found out is that so first of all in put privacy um in is tricky with zkl uh it's usually sort of a guarantee when you do Zer knowledge proof right you put in an input signal you label it as private and Bal boom you got a circuit along with the proof that actually seals the input um it's a little bit tricky with machine learning because ultimately machine learning is a program that extract the statistics of your input and it's also tricky because when when we talk about input privacy typically means you distribute the approver to client side right so I hand you the approver so that you can generate proof on your uh on your device and you just hand over the proof to me um so that also means that you're handing over the model essentially and the problem with handing over the model to client side is that then they get to run inferences over and over and over again for free and in that sense they can actually then train a model to inverse compute uh whatever input you have um so let's say right so I can basically generate billions of inputs and then run this through a model and then using these input output pairs trying to compute like an inverse uh model that will inverse your input um of course this is only limited to models that probably don't compress the dimensions as much right so of course like if you have a model mod that takes a picture and output like a teeny a tiny label right like an integer label that's very hard to reconstruct but there is a lot more other models that actually would not reduce that Dimensions as uh so much right and there's a probability of uh a possibility of people reconstructing that um so uh and we have actually done that with uh with in the uh archive paper that uh will be linked in the link tree later and so inside there is actually an experim where we actually try to use um ZK and then right so we just have like the input and output of a CK circuit and trying to reconstruct a model that can reconstruct all these inputs so that's um a bummer right so okay well maybe we not have we don't have input privacy always uh but certainly we can have model privacy right so it actually also suffers from the same problem which is again machine learning is extracting the statistics uh of these data and and what you would uh what happened with model privacy is first of all you have a limitation that it's a single Pro right it's a single Pro assumption so okay because Distributing the model is dangerous uh as I mentioned in the previous slide so okay let's take it back we're going to do like a single Pro assumption which is whoever holds the model will compute all the proofs right then the model should be secure because only theover has all the model weights um however uh actually if there is enough inferences that are made uh made in the uh from this model right and you get enough pairs of input and output you can still then reconstruct or train a so-called like euristic model that mimic what your model actually does right so even then ckb doesn't guarantee model privacy automatically you still have to sort of limit the inferences that is made by your model by either economics or scarcity so essentially limiting how long you used the same model for and before people can break your model then switch to the other model right so the bottom line what I wanted to show in you know both the input privacy and the model privacy case is that CK is a little bit tricky with machine learning and all this privacy guarantee doesn't come automatically it still needs to be aligned with economics and so then a interesting Paradigm comes uh actually the aura team encountered this on an if research uh Forum post which is someone proposing something called optimistic machine learning and he quickly joined the team and we have been developing this ever since so optimistic machine learning is the optimistic like row up of ckl essentially we're guaranteeing the security with Game Theory right so you have a submitter well you have a bunch of notes that will compete to be the first to submit a machine learning inference results and then the other notes if they actually want to challenge the result they can uh actually um uh they can initiate a dispute on chain and have an interactive thought proof so um in this case is definitely I guess a different trust assumption compared to CK which is you only need uh well for CK is pretty much trustless um for op you need like one at least one honest node out of the whole network right so the good thing about opml is you can have huge models okay here is just some example that we have already put on chain like stable diffusion and llama uh but it doesn't limit to that uh at one point we also have Gro which is like 100 billions parameters uh on chain it's just not performance as a model so we talking get down um so but um with ckl is only sort of practical right now with very well smaller models decision Forest gpt2 um the experiment we will show later also show that zkl is capable of doing llama but in the time scale of days or hours but the thing is well doesn't like looking at this graph and also thinking about with privacy in mind because that's one of the things that people challenge us the most is you lose any privacy right as as soon as you go op as soon as you go optimistic so hence then we think of why can't we combine sort of the best of both world right so we know that um with the previous slides then in ckl input privacy might be tricky model privacy maybe right we still need to limit the number of inferences that is made by that model so why don't we think about a way where you we can have model privacy or we need model privacy uh but then we also can run optimistically in some of the program so this is where we introduce opai um actually opai has a certain meaning in Japanese so if you're uh interested you can always Google what it means in Japanese but I think this opai is actually cooler uh so it stands for optimistic privacy preserving AI uh it's a really cool name for actually a very simple scheme uh which is well we run op on some uh part of the program and then we run ZK in the rest of the program okay so that's um actually this the the idea so we're going to um uh we're going to apply CK selectively in particular on the valuable parts of a model okay so this is something that um is more on the machiner side um have everyone heard of like Laura okay so Laura is the fine-tune version of stable diff well the fine tune fine tuning parts of the stable diffusion models such that it will fit certain style or certain kind of fine tune examples right so actually with for example Laura and on stable diffusion so stable diffusion is a open source model right at least stable diffusion 2 is a open source fully open source model so really there is no incentive to hide it whatsoever it can be run perfectly with um o opml but let's say if I have Fon a model well now those fine tun weights become valuable to me and in order to hide those weights actually I only need to ckf those weights that are important to me those weights that have been altered by Laura all right so this is actually the idea of opai which is we can apply ckl selectively and and it's great for fine tun models um it turns out that when I was thinking about this idea I thought stable diffusion would be a great example to implement this turns out that even if I just apply CK on the LOD Change Model weights of stable diffusion it's going to still take hundreds of terabytes of memory so this is how big machine learning models is we're talking about especially if we try to fit in CK right so even if we do it with stable diffusion and uh and just on the low out weights it's actually still going to take a lot of memories and be very timec consuming so we found a better example recently uh this is actually not in the preprint the first version of the preprint uh which is actually um not a fine-tune stable diffusion but a fine-tune llama so a fine-tune llama 7 billion in particular and also not Laura because Laura usually means you're going to change a lot of the weights in the Transformer part which M makes like making all the part CK very expensive uh but then with something called uh ref uh um right left actually the long term is uh low ref so low rank uh ref um you're going to basically do interventions in these like large language model and the interventions are actually just a matrix multipli oh Matrix projections so it's actually very small and in a llama 7 billion model there is 32 interventions um combining there are two million parameters so each of these interventions and the thing is each of these inter SoCal interventions can be coded in CK circuit separately right so you're going to have 32 CK circuits that essentially you can proof parallel verify par parallel and they are much much cheaper uh to prove and and verify compared to trying to do the whole llama 7B end to end so um yeah so this is what we did um I I actually just did this like last week so I did a very naive implementation in circum with snjs and grow 16 and actually just on my laptop so on my MacBook Pro with running M3 Max and here's some result compared to um two other prover that try to the whole llama 7B uh end to end so the black bars a is a recent uh result from legatron so they did uh they did it on a very high memory machine and they got result which is you need around 10 gigabyte of Rams in around 14 days and uh yeah right to prove it so the the time is actually in logarithmic scale um on the right side so um and and we're talking about when we're talking about this Benchmark we're talking about proving one token okay it's not one sentence it's just a single token um and then there's a more recent paper right after legatron has um distributed their results called CK llm I believe from waterl and they used a a100 GPU and was able to do it with in 14 hours okay um and so then our um very naive implementation actually were able to prove it in 2 gigabyte memory so pretty much any machine that you have on hand right any laptop um and then if we prove it all sequentially right just to make the memory uh Benchmark fair so if you just want two gigabytes of RAM but then you're willing to prove everything sequentially it takes around one minute and then you can verify everything parallely in very uh in a very uh short time because there a grow 16 proof um with quite uh with like minimal um proof size so this is um the idea of uh opai which is you can think of all these other models that we would love to put in uh what to have some ZK property uh but then it's also more attainable because it's like basically ready for a very large model okay so uh I'm going to to switch gear a little bit use the last few minutes to talk about okay let's say now so now we have made proving sort of real life model a reality or verifiable ml reality and usually the questions that we get asked the most in presentation is well why do we need AI for blockchain why do we need AI on uh on chain and I would challenge you to think about the flip side of the question which is actually maybe AI on blockchain is not for us but for AI um so just to see like how we can um use opml to sort of change some of the Paradigm so actually when we talk to AI companies they typically face two problem uh first of all of course their biggest competitor is always open Ai No matter what kind of models they do when open AI uh start doing their types of models right for example text to video then their whole startup is over um so one of the uh um things that they want to challenge right or they want to at least prove the customer is that hey we want to prove to you that we're always using like the same model we're unlike open AI which you know they have a slow drift of their GPT performance uh we want to show you that there is some verifiability we are guaranteeing some surface that we're giving you so that's a verifiability um and now on the other hand there is another problem which is so a lot of startups they also opt for open sourcing their models right so to gain a little bit of more market share for some uh for more users to use them but then when they want to start monetizing it then it becomes another problem right which is as soon as it's open source then it's very hard to monetize so um actually what we have done is we're have basically offer modeled as like a tokenized version of them so we're going to uh take model ownership which is every model is breaken into like tokenization and then we're going to pair it up with what we call inference asset which is inference results that have been done by verifiable Machine learning all right so you might think why would that be valuable well uh prior to having verifiable ml actually people would uh have already we already have ai art in nft right like we have always had that but the problem is no one know what model it has been used no one even know whether it's a original model from that artist maybe that artist just used some random model uh that they don't aren't even licensed to use right so um actually ERC 707 is the first AI based ERC that talks about while taking either ckl or opml and let's make this an endtoend process so now creatorship of like aigc nfts or AI generated nfts is a end to process of hey I proved to you that I had some input that have put through this model and here's the output right so now this becomes a avilable piece of data because all that also means that uh essentially if someone wants to make for example the same inference they don't actually need to do it again right so this is um some sort of uh it's a piece of asset that will be useful right so um yeah so uh that's how sort of like Aura has put AI on chain so I've talked mostly about the first part here which is like the optimistic machine learning and to realize optimistic machine learning uh we have actually offered like it as a back end of our onch AI Oracle so our onch AI Oracle is where developers can um call a uh model model inference directly from the smart contract and then our uh opml Network would actually Supply the results back to the uh Oracle and then through like a call back to the user smart contract right and then uh the initial model offing is the other counterpart where we're showing that there's like value from verifiable machine learning uh through like tokenization cool so uh that's the end of my presentation um the both the opml and OPI is actually pre-prints on archive and it's linked in my link tree [Applause] do we have any questions to the audience yes hey can you hear me okay cool talk um how do you do verifiable ml if ml is non-deterministic and blockchains are [Music] deterministic yeah that's actually a very good question that we also get a lot um so it's contrary to public belief machine learning is actually pretty deterministic or at least can be controll to be deterministic through a lot of ways so um first of all if you're talking about the intrinsic randomness of the machine you can specify a random seat uh so that's the first point but then you still get probably Precision errors from like folding Point um arithmetic so one way to do so is through software float uh which is a uh a a feature that you can enable for a lot of um uh targets that is compiled from like go or rust or llvm um and then the second way is actually you can also enable deterministic computation on Cuda it's going to add like a 20 to 30% overhead of computation uh but it's at least like on the same hardware and the same OS like assuming everything is the same but on two like physically different Hardwares you can still uh do it deterministically so it's possible yeah does that answer your question yes we had one more question yes thank you for the talk uh I have two questions actually first um these inference assets that you talked about are they going to be produced by Aura and if yes are you using some crowdsource gpus for the task uh and the second one is um FH in confidentiality is a big talk right now in web3 and AI do you plan anything in that regard thank [Music] you right so uh with regard to the first question so we consider Aura to be like a middleware so uh Hardware agnostic um of course up to certain point from like the deterministic perspective but um so we are going to basically decentralize the node and for the note probably then they would be running on decentralized GPU or different Cloud surfaces so uh that's something that we at least hope to happen like we don't wish for all our notes to be on AWS or gcp uh so I hope that answered the first part of your question uh the second part of the question is about fhe uh I would say we are closely sort of like monitoring and looking at the technology but we also realized that to have reliable fhe you also or or at least like trustless fxg you also first need to do a like this zkf the the FH the the encryption process right so you can't really have FH without CK being ver viable uh so I think that's also the biggest sort of bottleneck that we're when we're looking at this technology right it's first ZK has to be quick before we can do uh use FH we have time yes but we have time for a quick question so uh thank you it was a great presentation for the optimistic version does it mean that every time a user wants to make an inference they send all the information about the inference on chck um and what's the impact on on privacy for that part right so uh exactly so at least for the C uh well yeah so for our current implementation everyone has to send their full input and uh the models are actually not any customized models they are open source models that is like already in our offering um so in in in short there's no privacy and that's why we're exploring uh but yes so but then then also even even when we're exploring like combining CK part it seems like it's easier to protect the model provider more than to protect the the input provider uh because of the limitation of ml yeah we and that's a wrap please once again give it up for caty

Automatic transcript — names and jargon may be misspelled.