New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

CleanUnicorn - Beyond "Trust me, bro" Engineering LLM reliability

ETHCluj MeetupTue, Oct 7, 2025, 12:00 AM

Large Language Models are rapidly becoming part of our everyday life. However, this requires an immense increase in computation, handling private data, shortly put: trust. This talk delves into the critical question of how to build networks that run LLMs for users in a way that increases user trust. How can you know that the remote GPU is actually running the LLM that you expect it to run? This talk presents a way to verify LLM compute, whether that is training or inference. The goal is to present a practical technique that can be applied to decentralized networks to enable trusted LLM compute.

Transcript

And I want to kind of um first tell you a little bit about myself. I'm a developer. I did a lot of security audits in the in the blockchain space. Um I'm also an investor because I'm working with Eden Block right now. Uh which is I'm a technical partner there.

Uh and these are some of the projects I worked on and um I am a tech nerd. Yeah. And you can follow me on Twitter if you want. So um the the important part is like I'm want to make clear what you're going to get out of this talk. the um if you stay until the end you can understand a technique or at least a kind of a concept um by which you can verify correct execution of LLMs but all of this onchain which is like you're going to see it's a big problem so let's first kind of define the problem um I I it's not about privacy so it's not about data collection retention you know deleting data or encrypting it or hiding it in in any kind of way.

Um I'm like but however if you're interested in that I had a different talk at the cipher pong congress devcon last year uh like the privacy paradox in AI. I kind of tried to explain that the more data you give to the AI the better responses you're going to get back. So that's another important thing to do. And what I'm talking about right now is mostly correct execution. So we want to make sure that the u inputs that we are sending to the model uh our prompts our you know user prompts our system prompts and everything the parameters um and also the models that we're expecting things to run are the ones that we expect like also you know the model version um the model parameters like everything should be correct because kind of every time we're doing that we're paying for that execution.

So why not get that? Um, so this is an Ethereum conference, but everybody's talking about AI. You should try this experiment. Like every time you hear AI do one push-up, let's see what happens. If you think that's a bit too much, uh, maybe when you hear AI agents do one push-up, you're still going to get really buff.

Um, like I'm really really interested to kind of understand like who uses AI here. Like if you're using any kind of chat GPT like chatty interface. Okay, great. Keep your hands up because I want to understand who uses AI through an API. So if you're asking an API to do something for you, okay, that's that's a lot more people than I was expecting.

And who's using AI run locally? If you run your AI on your own machine, quite a few people. Wow. Okay. Um, who trusts Chad GPT?

Okay, we have a few people.

Okay, so like when you send that request to um to your API, how do you know that they use that model with those parameters? They haven't truncated the model in any way. They haven't altered the output. They're not giving you a cached response. Like how how do you know that's that worked?

Like I'm literally asking you like does anyone have kind of an answer for this?

I'll pass you my mic.

It's like a bit tongue and cheek, but I think like 40 is especially like a psychopant. It it tries to please like please you a lot and I I actually can get a sense of like oh this is 40. This is like the guy who always tries to please me. But besides that I actually Yes, it's a good question. I don't know.

So it's a bit of a heristic like you're kind of yeah you're kind of you have learned how the model behaves and now you kind of trust that the response is from that model but like most of the time we don't really know and we trust the systems that they did the right job for us. I want to give you an example. This is I'm not sure if anyone heard of open router. Open router is this centralized place where people go send API requests and then open router routes your request to either like something like an open AI a deepseek or even open source models like llama 3 and you can either choose the specific model or they can route it based on some kind of algorithm and in this case you have you're innocent user you make the request to open router and then you trust that they route the model to the correct um model runner for you. So you have you put trust in them and then they also put trust in the model runner.

So there's even more um trust assumptions in that case and I've ran some some experiment. I just opened up a Python notebook, created a bit of code and I'm asking Quen like who are you? And I asked it 10 times and he tells me I'm Claude. I'm JPT. My name is AI assistant.

And only one of them said my name is Quen. Which was like what I was expecting. So like what's happening here? Is the model wrong? Are they lying to me and they're not running this on Quen?

I don't really know. maybe they also don't really know. Um so that's one way to kind of cheat the system. Uh a different way is through quantization. When you have uh quantized model um you're kind of taking this very smooth model that you see on the left and you're making it a bit rougher.

You're compressing some of that data into a smaller um space. What and thus you get lower precision. What you're actually doing in that case is like let's say say that you have that model on the left you have very very large precision you have lots of decimals and then you're just cutting off some decimals and you might say oh maybe this doesn't is not that significant but if you look at the difference in size of models this is the same model so this is deepseek R1 32 billion parameters um FP16 which is by the way not the largest one um is 66 GB. When you quantize it to Q4, it gets to 20 GB. So, it's like a third of the size.

That means you need cheaper hardware to run it. You don't expend as much energy. It's much much easier for the model runner to say that yeah, I am running DC R1 32 billion, but it's not the one that I was expecting. It doesn't have the same uh power. So in this kind of examples, we learn to cheat in a few ways.

There are many more ways to cheat when you're running models, but still kind of say that you're doing the right thing or you're providing the right service. So there are some ways to provide solutions for this. We're going to explore a few. Um maybe the most obvious one is cryptographic proofs. Um okay, so you have cryptographic proofs.

You take some kind of message, you sign it, and you give the person like, "Yes, this is the signature. I signed it. It's totally fine." The thing is, when you're signing a message, you're not signing an execution. You're only signing, "Oh, this is me.

I want to run these uh this prompt and I'm expecting some kind of output, but I don't know what will be executed. I don't know what the output will be. I'm just signing my intention that I want something done." So how do you u sign an execution? How do you prove an execution?

It's zero knowledge proofs. Um the thing with zero knowledge proofs is they are really really expensive. Like one does not simply build a zero knowledge proof. Um and that's because you have to transform all of that code to math. So all of the execution paths have to be coded into a polomial and then you have to generate a proof on top of that it's really complicated and if you take something like uh a token and you uh generate one token with a model like for example using GPT3 it takes less than one second to generate one token.

one token would be a word. Doing the same kind of work and prove that you generated the model through GPT3 takes 13 minutes. And this is not a sentence. This is oh sorry this is actually a sentence. So you would have to uh wait 13 minutes to get a sentence back and the proof that the execution was correct.

So it's totally unfeasible. Um but what about trusted execution environments? They might be a solution. um like who who knows what's uh TE okay a few people like for for everyone I see trusted execution environments as kind of hardware models you have some kind of hardware which is isolated from anything else maybe you have a hardware key there or a private key and you have some kind of execution mechanism execution possibility so once you send code to that it does something you get a response back and you're pretty sure that the execution was correct um these uh hardware wallets which they're not technically TEES they do a very very simple computation. So they're very useful when you want to sign a transaction but they cannot run models for you.

They're just not powerful enough. Um and also you kind of have to trust the hardware. um you put your trust in the manufacturer, you put your trust in the person who put the uh private key in there or you generate a new one. So there's there's a lot of another type of assumptions there. But maybe one of the worst problems is this is specialized hardware.

So something like this already exists. Um the most popular one is Nvidia H100 and it costs at least $25,000 to to buy that and it's pretty good. It it does its job. It can run models for you. But if you look at this table, uh you're you're going to see Nvidia a lot.

So you have to you're locking all of your systems into one hardware producer. Um Nvidia or sorry the uh Intel SGX is also a trusted execution environment, but it's not really right for models. So Nvidia is pretty much the only producer which can solve that problem for you. Okay, different different idea. social consensus.

Um, so people who are working with Dowos or just do consensus in general know that it's a very difficult uh job to coordinate everyone and to convince everyone that okay what we're doing is right. we're making the right decisions and not just that it's like we are um there's no um sorry and there's like what you're doing is right and we're making the right decisions but in this case it's about execution so you kind of have to ask multiple people to do the same kind of compute and give you the answer and verify the uh the check sum of that and you say okay that's fine however you still ask multiple uh hardwares to do the same kind of compute and even if you see a divergence in responses, you cannot know who the right person is without doing the computation one more time or you maybe just take the uh the majority but you still cannot have a smart contract that can kind of prove like who is in the wrong and who is in the right. So there is an alternative solution to this and and this is already exists in the wild. it has been implemented. Um but in order to get there we kind of have to code the um understand how inference works and I'm always when I'm when I'm doing doing these talks I'm just uh trying to simplify things quite a lot.

So there are a lot of details which are not explained here but the concept is there. So let's say that we want to generate something based on like generate an output based on the initial input of hello. What you can do is you take all of the the model, you encode that into a series of operations. So you're taking um that input and taking it through the model and then you get one additional word back. So you take hello, do all of the operation and you get hello world back.

Okay, good. So now you kind of understand the inference process in a very very simple way. Uh in order to generate one more word, you have to take hello world do all of that once again and you get another word. The cool thing is if you if you look at this you can create a merkel tree based on that. So um I guess a lot of people are pretty familiar with Merkel trees.

We've we kind of understand how how airdrops work. Um and yeah actually how who knows how kind of a Merkel tree works. You don't have to implement it but principles. Good good good. Uh the idea is that a Merkel tree gives you some kind of assurance that the data you um encoded is um the one that you expect.

In this case it's kind of a very expensive check sum because you have to do a lot of steps to get there. But it has some very interesting properties. So you ask the people who do the computation to generate this Merkel tree on top of the inference and then you kind of compress all of that into one uh route. Once you have that you can kind of start navigating the tree and uh this is where things become very interesting. So let's say that for example you give uh the same kind of operation to two node runners.

Each of them generates a tree. You only check the root of the tree. If they match, you're pretty sure that they did the same kind of work. And not just that, but maybe they're right because they um maybe didn't want to collude or it was difficult for them to collude and and generate the same tree for you. Cool.

But what do you do if they don't match? So in this case we have two trees the roots don't match but you have to find out who is in the wrong like who did the wrong computation. What you can do then is start traversing the tree from the root. So like from top to bottom. So you check the root you see okay it doesn't match and you ask both of the runners give me the children of these roots and then you start uh following the children who are different for each of them.

So in this case you first go to the left children. You then uh ask for the the children again. You see that the left one matches. So that means okay everything that's to the left is fine. They did the same kind of they had the same kind of output.

Let's traverse on the right. So you go right you go left until you reach the leaf. So boom. Echo. We found out where they their uh computation started to diverge.

That's cool. But still who is in the wrong here? um it's it's just one computation. It's not the whole tree. So you feed this operation to a smart contract.

You find out who did the computation right, who was incorrect, and you have strong proofs of um who was uh who was correct or who was incorrect. What you do then is, you know, punish the the person who did the wrong computation and reward the other one and good you you proved LLM computation on chain in a scalable way. And this actually works. Um I actually implemented this for for a company there. This is exists in the wild and uh we're going to see a lot more alterations of this kind of algorithm in the future.

I think it's a it's a good approach. It's not the final version of what we're going to see, but having blockchains which verify computation uh for AI, it's going to be it's going to exist more and more like blockchains are extremely good tool to coordinate to incentivize and to punish people. And AI is probably the um AI is probably the the uh like iPhone moment of of blockchain. So these are a few uh references. Um I'm going to I'm going to share these slides on Twitter.

The first one is uh Verde. Verde is actually the um implementation that uh I implemented which is based on Agatha. Agatha was launched in 2021 as as a paper. Uh Verde was launched this year in February. So it took quite a bit of years to distill that initial idea into something that works practically on chain.

And some of the others are kind of uh expressing like what it means to do verification on chain. How difficult is it whether you're doing ZK or other kinds of things. So to kind of sum everything up, uh ZK is very very heavy. It doesn't really work on chain. Also, it doesn't really work because the set of computation is extremely large.

Uh blind trust in in hardware or in social consensus. I think it's too risky. Uh pure social consensus is very weak. You don't really know who is in the wrong and there's no way to prove that. Uh but I think hybrid approaches can scale.

There's something there that can be developed in the future.

Let's go to the questions. What practical advice would you give to a non-technical person who uh how can they ensure that their AI models are not hallucinating? Okay, cool. Um, not very related to the talk, but I can try to answer that. Uh, you have to learn how to ask questions.

Like, for example, when you do some kind of research, you have to ask non-biased questions. So, when you ask your AI to do something, it will give you an answer. Take that answer, put it into a different AI or into a different thread and ask it to be critical of that answer. Like, play the devil's advocate. this might help.

Um, and this is maybe the fastest way of doing that. The other way is you have to check everything. So, it's it's not super super scalable. AIS will still hallucinate. It's it's getting much much better since a few years ago.

Uh, but it's it's good not to trust them completely.

The next question is, do you really need a smart contract? Oh, it moved. Uh, do you really need a smart contract for the verification or could you just run the comparison on your own infra?

Yes. Okay. So, you do need a smart contract verification. If you run the comparison on your own infra, you're basically have to run the whole execution from beginning to end. And if you're doing that on your own infra, why would you like you're it doesn't make sense to ask anyone else to do that.

Just do it yourself and get the output. Awesome presentation, awesome answer. Why do you think we need AI on chain and what's the real purpose for having AI on chain

in normal internet? We have a few companies that control a a huge amount of what goes over the wire right now. Um I I don't think that's where we want to be, but we kind of got to this place right now. We we're already seeing very large and some of the same players who are building AI models. Um and we kind of need an alternative power to at least put them in check and make sure that they're they don't have complete control over this type of technology.

So I think that decentralized technologies can be a strong power that fights with the likes of OpenAI, Meta and and many many others. Can I choose one question because it was kind of interesting.

Yeah, you can choose your questions.

Um, okay. So like the last one is uh can't zk proofs use a hash of an output instead of verifying uh token by token can we get certainty of origin that way anything more than uh we are after so um there's one thing I haven't mentioned but when you post your root on chain there is you have to do a few more things there you have to do a commit reveal you can't just post the hash of the tree because let's say that I know that me and you are working on this. I'm waiting for you to post the hash. I'm just copying the hash and I'm posting that myself and the contract will say, "Oh, it's verified. They're the same."

So, what you actually have to do, you have to do a commit reveal. And that's how you you hide the uh the hash of of the whole complete output. Um you don't need ZK to hide that hash. What you would want ZK for is the for the complete execution. We already have a way to hide the hash.

We don't have a way to verify execution.

Yeah, let's just take some questions from the audience. We'll start from here. Think about your questions.

Just a basic question um about a problem itself. So you chat claude have a bajillion daily active users. if something in their execution pipeline goes wrong, presumably many people will figure it out. So the chance that you are the first one that interacts with the model when something goes wrong is very very like small. Um so I'm like just like trying to understand the problem in the big picture.

Is it more about um like biases like political biases that models can have or governments getting in the way of like the actual model being executed? what is the big problem in terms of you know because in practice you would like the the the community would figure it out if something is wrong with the models

right

uh very very valid question like why do you care about this in principle

open AI can indeed like you're going to see that okay maybe the model is just became a bit stupider but I'm not sure right now um the idea is that open AI and all of these companies can swap models for maybe a bit more biased models or updated or truncated or anything like that and we've seen that with chat GPT especially in the beginning when you couldn't choose the model and they would choose the model for you they would update it update it behind the interface but this is more okay for a as a consumer maybe you don't care so much about that if you're a company and you want a specific model to be used maybe that model has been trained for a very very specific use that you need and you either have um medical data or banking or anything like that. You want that model. You have to be sure that what you're paying for and what you're expecting to be executed uh is is exactly that thing. And sometimes it's your your own model. So you don't want people to quantize your own model and then use it because I I I built my own model.

I want to use my own model, but I don't necessarily have to run all of the infrastructure myself. So you're kind of um segregating you're separating all of these operations into something which is more manageable and you're finding the right players to execute the model, train the model, do anything like that. It's going to be a bit a lot more diversified in the future.

Sorry. Sorry. Can you repeat the question because we're recording it?

Question. I was just just saying that you can also use the centralized computing networks. So this is exactly what it's doing. It's using decentralized computer networks to compute something for you in a decentralized way. Do we have any more questions?

Okay, think about your questions. I'll read one more question for you. Uh how do you think about multi- aent consensus and multi- aenic coms verification will evolve in the short term? Uhhuh. Multi-agent consensus, multi-agentic verification.

Uh this is um mostly exo. Okay. It's an AI agent question. So we have to do all all of us have to do one push-up for this. Um this will exist in the future.

We're going to see different types of models that verify other models. Um it's imminent. it's already happening. If you're looking at even if you're coding with CL cloud code, um you can think of it as different agents are doing different things. Maybe you're running something, it's taking the output of that.

There could be different agents. One thing I really when I build an AI agent swarm, I like to have very very small very uh particular agents for a particular task because that's where they uh achieve their task be best. So for example, I'm asking one AI agent to write something for me and then I'm asking a different one to change the style of that writing to a to a particular thing. I'm not asking one agent to do multiple things at the same time because it can forget. Sometimes it does, sometimes it it doesn't.

Uh and it's a lot faster and cheaper to use smaller um models to ask for specific types of things.

Okay, number four. Question number four. Here's one silly question for you. I'm just joking. No smiles.

No. Sorry. The joke didn't land just because sending it might be a silly question. I'm not making any better for myself.

Um, so the question is um maybe a basic Oh,

where do you go?

So the different models.

Yeah, that question.

Do you want me to read it or do you want

I can read it. So the different models we run the computation and you reward the one that does it. Correct. But isn't there a lot of computation lost might be a silly question. Um you're you don't really have uh models for the computation.

You have machines let's say GPUs that are doing this computation and there yes you have models on top and and you're doing a lot of that. There is this um performance hit because you have to build the Merkel tree. have to do some additional steps on top and you kind of have two machines that are doing the same kind of computation for you. So you you double the cost immediately because you need the consensus. Um so yes it's it's not as efficient but I don't know exactly what the numbers are.

A lot of people are saying that AWS and Google Cloud Platform have high markups on their hardware. So ideally let's let's create some competition for them. Let's improve the algorithms and let's find um let's let's put them in in their place because they shouldn't be the only ones who build models and run models for us.

Do you want to keep answering questions?

I can do it. Yeah.

Okay. Um um what did you Sorry. Whoa. Whoa. Okay.

Okay. Let's let's turn this into a discussion because the questions are consistently moving. But um who wants to ask a question on the microphone? How the hell? All of you are writing questions but no one Okay, thank you.

I was like, what's going on here? There are AI agents in the room. So I think the part that's most confusing to me is like has were those papers like especially the one that's verifying the the like process through which the model goes when you ask hello that it says world is that consistent regardless which like what time of day which part of the model it hits how much stress the model has because that inference is running on the model cons like um at the the same time like you're not the only person using the model that's the idea right behind the kind of shared model model interface

so is there any risk of like actually the hash being different because it took a different neural pathway because like it's

DNN right so it'll be neural network layers

yeah so so this kind of question has is is very very valid and it it kind of shows that you understand a lot more about how uh models work and how the hardware works Um there are a few things that you can control when you um do an some kind of inference. Almost always you add some kind of random seed to um to the input. So you can you can also control that random seed. But even if you do that with having the same random seed, you could still have different outputs. And this is sometimes because of the hardware.

uh Nvidia has some kind of optimizations and they maybe they do the operations in a like in a tiny different way and then all of that error gets compounded and you get a different token on

well and it's not even the hardware I think the hardware is kind of the least denominator here from an error propagation perspective like is the nature of the neural network and the way we design these LLMs as probabilistic systems that will kind of have a 90ish% 95% adherance to a method and then there's also like a part that will just not make sense.

Okay, so if we consider that hardware is not an issue, uh we still have very concrete numbers in the model. So they don't change and if you have exactly the same input, you're always going to have the same output if hardware is not an issue. I I I I disagree here because I understand it's probabilistic, but the probabilistic part comes from the random seed. And I'm saying that you can also control the random seed. So if you control the random seed, it's not random anymore.

It's not probabilistic. It's completely deterministic in that

Automatic transcript — names and jargon may be misspelled.