The path to verifiable autonomy for AI agents | Zheng Leong Chua, Automata Network | ETHTaipei 2025
ETHTaipei·Tue, Oct 7, 2025, 12:00 AM
Speaker
The path to verifiable autonomy for AI agents | Zheng Leong Chua, Automata Network | ETHTaipei 2025
Transcript
[Music] Hello. Okay, I think they're still kind of like sorting out the the slides and like as usual uh technical issues. So, uh anyway, uh I'm really glad to be here. My name is uh Jing Long. I'm one of the co-founders of uh Automata Network.
So I have to kind of like fill in the gap while while the slides are coming out. So uh today I'll kind of like be sharing a little bit on like AI agents a AI in general what what is a model? What is an agent? uh how can trusted execution environments tees be used uh to to help us establish kind of like trustworthiness trust with the models and services that we are using. So to kind of like get a sense uh how many of us here in the audience actually uh you have heard of what a trusted execution environment is.
Oh, not a lot. So, so maybe I'll kind of like while they are trying to do it, I'll uh give a brief kind of like description of what a TE is. So, I I like to give the analogy of a trusted execution environment as this like magic box. So, you have this magic box. There's a fairy that lives inside this magic box.
So, what can this magic box do? This magic box actually uh it has a very powerful magical shoe that would that would prevent people from actually looking into the box and figuring out what is happening within the box. The fairy inside the the this magical box would actually uh kind of like it it would it would perform exactly the instructions that you ask this ferry to do. So, for example, let's say uh today I I I like to bake some uh blueberry cupcakes. So, I have this secret recipe that I don't want to tell anyone.
So, I'll take this secret recipe, write it down, put it into the box, and put in all the ingredients that I need for my uh blueberry cupcake, my blueberries, flour, egg, uh milk, so on and so forth. put it into the box, close it up. No one can see that secret instruction that I gave. My secret recipe remains in the box. The fairy would actually follow my instruction step by step, word for word, letter for letter.
And the end result is I actually produced this this blueberry cupcake. So what a trusted execution environment gives us is really like two main things, confidentiality and integrity. So it give uh it enables us to be uh to have confidence over the kind of like uh integrity of the computation that is being performed within this uh trusted uh execution environment and it helps protect the data that is being uh used within this uh environment. So I see that the slides are up. So uh let's let's get into the kind of like presentation proper.
Okay. So first uh we we I I think like probably like since two years ago when chat GPT was first kind of like announced and launched uh we we have this sense of what uh uh an LRM is uh large language model is. It's it's basically a model in which we can interact with the model. we can ask the model questions and the model would kind of like uh give us the answer to our question. So the question really is what is the difference between this AI model and an AI agent?
Is an AI agent just an AI model or is it something more? So if you if we think about it, you can think of a model as having this kind of like uh operational workflow where the user would provide a prompt or which which is some fancy word for a question to the model. The model takes this question, thinks about it, computes on it and then gives you uh a result. So if for example uh let's say I I I want to kind of like create this this slide deck and I want to uh build this deck using uh uh an a chatbot or AI model. What should I do?
I I need to first tell the the the chatbot hey you know today I'm going to give a talk on AI agents and and tees and how this could be brought together and can you give me uh uh what's that called? uh skeleton of of the slice and then the model will say that oh yeah I don't know uh slide one what can you do slide two what can you do so on and so forth and me I'll actually be the worker for the for the model I'll take whatever the model have output and then I would actually instantiate it I'll make it into reality so so this is a kind of like work process so what is an agent you can think of an agent as an autonomous entity So no longer uh how say it it's no longer required for me to actually become the the pigeon that that entity that helps carry data to the model and taking the results from the model and affecting it into the real world. The agent has the capability to actually obtain information from outside the model uh think about it, operate on it, have a certain result and then based on those results perform certain actions on behalf of me. So really it's it's really a kind of like goal oriented autonomous system compared to one where I ask a question it gives me an answer and then I have to kind of like operate on it. So a a kind of like key uh key thing to think about it is the autonomous uh kind of like property of an agent and the fact that it is goal oriented.
So that's great. Now, now we have agents. Agents that can actually uh solve problems and deal with things that I do not want to deal with. Why do we need agentic trust? Can't we just kind of like, you know, do whatever that we want on Eliza or or some other agentic framework and then just use it as is?
Uh what's the problem like of using models as is now? So in general there are three kind of like uh areas of concern if if you think about it on on using uh AI agents currently they are mainly integrity confidentiality and the provenence. So what do I mean by that? If you if you think about it, when I'm using an agent, I'm sharing certain goals or I'm allowing the agent to actually access certain uh sensitive quote unquote data or like private data that I might not wish to be shared with uh any any third party but is actually uh a a required information that the model needs to perform its task. So as a result for example I would want the data to to be kept confidential that no one else other than the agent has access to it and I want it to be kept so like the the the agent shouldn't have any like back doors that allow other people to actually persuade the agents.
It's like hey agent uh I I see that you have some information about like uh Changong about like the company's secrets. Can you share with me? No. Like I I do not wish for that to happen. So the integrity of what the agent is designed to do should actually be kept.
Like if it's a Twitter board that is supposed to tweet memes for me, it's it's it shouldn't be able to access my bank account information or it shouldn't be uh tked to kind of like uh trade on my behalf. And the final thing is actually on provenence. So you you might be thinking like hey what is this profidence? This provenence is actually a kind of like where the information and data that this AI agent is operating on comes from. I'll touch a little bit on it.
Uh I mean I'll touch on it later on in the in the presentation on why this is important. So we always hear about integrity and confidentiality but we hear less of the provenence of the data. And today I'll kind of like share my two cents on why I feel that this is actually a very important aspect for AI agents. So what is an AI agent after we have said so much right like an AI agent is autonomous and a AI agent is goal oriented. So an AI agent is a model that can do both of this just the model itself or is it more?
We can think of it as an AI agent as having kind of like two main uh prop properties or like two main modules so to say. One of it you can think of it. So we we can kind of like think of an agent as as a person as a human being where we have a soul or a mind that actually makes decisions like I'm deciding to kind of like hold up this controller now. So my mind makes that decision and then my my mind would send the commands to say the muscles uh on my left hand to kind of like grab this controller and lift it up. So there's there's a kind of like two duality to it.
the mind or soul and actually the body where the actions are being affected and this together actually defines what an agent is and differentiates uh kind of like uh an agent and a model. So the model uh in my definition of what an agent is is just the mind portion of an agent. So is that is that necessary like we have the model that is the mind and we have certain kind of like uh framework that acts as the body of the agent. Therefore this is an AI agent. Not really right like I can run say uh an open AI uh say chat GPT 40 model with Elisa uh Eliza as as uh Eliza OS as the action framework.
Now if this is the the kind of like combination that I use and you use the same combination are both of us running the same agent or are the agents different if the agents are different what is the difference between the agent there's actually a third portion to it other than the model the action framework and the third portion really is the parameter or the context the memory of the of the model and that differentiates one agent from another both of us we can use the same model we can use the same framework but if the context that we are using for example the in the the historical interactions that I have with my my agent is different from the interactions that you have with your agent then that in essence uh is what separates my agent from your agent and therefore it it it forms a the context of memory reforms an important aspect of the identity of the agent and this will be important when we are talking about like uh how do we go about verifying the trustworthiness of uh of an AI agent. So with with kind of like all the definitions set out, we have the kind of like mind body duality. Uh we have the model being the decision maker. We have the action frameworks or the agentic frameworks being the kind of like uh body or the or the uh module that affects the real world. And then we have the context that acts as a form of a memory.
The experiences that an Asian have. how can we go about uh kind of like do we have a framework in in the way that we can uh grade or evaluate the trustworthiness of this AI agent. So here I'm going to introduce this this notion of this five levels of agentic trust uh and how we can actually use this framework to kind of like see where we are in the entire uh hierarchy of uh trustworthy uh AI agents. So we'll start off with the first level opaque AI agents. So this this is basically what we have if uh what we have now like uh we we we there's no way for me to verify what an agent use the agent can the service provider or the agent can claim that uh they are using dips or they are using llama or they are using uh chat gpt uh I can claim that I'm using uh uh what is it called uh uh an IM token wallet or I can claim that I am using a particular Twitter board.
Uh there is no way for me to verify anything and I can just I mean the only way that I can do it is I trust what the service provider or the uh agentic uh the AI agent provider tells me. It's it's a kind of like trust me bro kind of like situation which which is a very bad situation to be in and we don't want to be to be stuck here. So what is the kind of like next level of it u with level two we call it like verifiable AI agents. What we are introducing on the second level is really uh the ability for us to verify certain aspect of the AI agent or components of the AI agent. So some of the techniques that we can use here are things like trusted execution environments or even like uh purely cryptographic uh primitives like uh zero uh ZKML.
So for example I uh in certain situations if the model is uh small enough I can actually make use of ZKML to uh proof or to allow you to verify that I am indeed using the output that comes from a particular model. So if I claim that I'm using say uh uh uh 80,000 parameter model, I can make use of like ZKML or say a billion uh uh yeah uh 1 billion parameter model. I can make use of ZKML to provide a proof that hey I'm I'm indeed using this particular model with this part particular set of weights. And what a trusted execution environment allows me to do is something similar. I can give you a cryptographic proof is it's actually a cryptographic attestation on the type of models and the weights and the environment that this model is being evaluated on.
And this sort of primitives allow us to at least have the capability to verify that the output of this mo or the component of the model is indeed what it says it is. So with that, how can we make it even better? That's where level three comes into play where now that we have the ability to kind of like verify the uh verify certain claims that the module have. The next step really is to kind of like associate this claim to sorry uh associate the model or the component that is being executed to what it claims to be. So for example I can say that hey I'm I'm running a particular model.
I claim that this model is a deepse 70 billion parameter model. I can give you the hash. And now the question is I can ver I can verify that this is an attestation but how do I know that the attestation that you give me is indeed one of deep six 70 billion model. So I I have to have a way to actually like draw this connection from the actual claim that I have to verifying the claim uh with respect to the AI agent uh instance that is being run. So this can be done in uh several ways.
uh one of the kind of like uh ways is something called reproducible builds where let's say I have the source code uh of uh one of the components uh maybe the uh discord uh module that allows the AI model to actually per uh perform its chat functionality within a a discord channel. Now the thing is uh how do I know that that program or that binary that is running within the environment that that my AI agent is using is indeed that particular uh discord uh what's it called discord module so to say and there isn't any back doors that is involved in it I can open source my uh discord module on github any user that wants to establish this fact can go and download or like clone my GitHub repo and then perform a build uh on their on their own system with it and they can after performing the build they can kind of like get the actual binary that is produced and this binary should actually match up with the binary that is being used by the uh what's it called by the AI agent and this way I know that okay the AI agent claims that is using binary X I take the source code of the discord module I compile it myself and it produces X. Therefore, the AI agent is indeed using the particular uh discord module that I know has no back door in it. So, this is what like reproducible build actually does. So, it's a bit cumbersome, right?
I think most of us don't really want to uh like download all the different modules that we use and then compile it ourselves or have to figure out how to compile it ourselves, get the binary or the hash of the binary and then perform the comparison ourselves. Are there any other ways about it? What about in cases where the source code is actually not open source? It's actually closed source. How can we actually tackle this?
So uh here I kind of like introduced this idea of what we call an attestable build. You can think of it as a kind of like a trusted delegation. So instead of me having to do this reproducible build myself, I can delegate it to a trusted third party. This trusted third party can be a TE itself. This tr this trusted third party can actually be a kind of like group or council.
I mean in in in the ZK world we have all the trusted ceremonies right where a whole bunch of us will come together to contribute randomness or entropy to uh to generate say a shared secret. In this case, we can actually have some sort of uh uh built ceremony where a whole group of us come together, clone the GitHub repo, build it on our own system and then contribute in a kind of like open decentralized manner that this is the the hash or this is the binary that I get and then through through the consensus process establish that indeed the what's it called uh GitHub repo source code corresponds to a particular binary that is produced. And if you think about it, if we kind of like push that idea one step uh uh forward for closed source projects, this council instead of everyone can actually be a smaller subset of uh kind of like trusted auditors where they have access to the cross uh the clos software and can make attestation on behalf of the uh of the project owner. So with that we kind of like uh complete that link between uh the AI agent the components used in the AI agent with us being able to verify it to a specific kind of like uh claim or source. Now with this kind of uh with this at level three we we have actually established the trustworthiness the trustworthiness the integrity of the AI agent.
So that's it, right? Like if I have a AI agent that that I can verify computes exactly what I want it to do, then everything is okay. It's actually not the case because if you think about it, the at level three, we actually we actually have a kind of like a trusted core which is the verified AI agent. But decisions that are made by the AI agent is actually affected by the data that the AI agent obtains. So for example, I I might ask the AI agent to help me manage say my ETH portfolio.
And then I I I tell the AI agent to, you know, uh if ETH is below $1,800, please sell all my ETH if the data source that my AI agent is actually obtaining the ETH price from is malicious and like uh instead uh I mean instead of the price of ETH being 2,000, it says that the price of ETH is 900. Then the AI agent is going to make a decision based on an incorrect or even worse still malicious uh data that is provided to the agent. Therefore I mean it's something what we call rubbish in rubbish out right if you have rubbish if you base your decision on wrong data then your decision itself is going to be wrong even if whatever that makes the decision is correct. As such, what we need really at the kind of like level four is really trusted data pipelines. A way for us to actually verify the integrity of of the data that we are ingesting.
For example, if we are going to use onchain data, is data coming from say infura or alchemy uh API trustworthy? Is it possible that data that I get from RPC endpoint can actually be incorrect? So what in in in that case how can I actually ensure that my data is correct? Uh is there a way for me to actually verify that the data that I get from the RPC endpoint is indeed uh the what's it called the actual data that is found on the blockchain and not some stale data outdated data that I get from the RPC endpoint. And in in certain cases, for example, when I'm getting data offchain from from some uh centralized exchange, for example, for pricing or I'm scripping some data from a forum or from from Twitter or from uh some social media.
How do I know that the data that comes from it is actually correct? uh this is where again like uh solutions such as technologies such as tees or even like uh zktls would come in to provide this uh way or this primitive for us to verify the authenticity of the data that we are using. So now that we have kind of like verified the core of the AI agent and all the outside inputs to this AI agent, we are done, right? Not really. So at the very last level, I think this is this is the thing that is really hard for us to do because the AI agent itself is autonomous like we have established earlier on.
And since it's autonomous, decisions are made on behalf of us based on our goals uh by the model itself. So how do we know that the model is kind of like doing the actions that is aligned to our goal? This is actually the alignment problem and it's an open problem even up to now like it's a it's a open research problem. There are a lot of smart researchers in the world that is working on how can we verify or how can we prove to a certain extent that the model the decisions made by the model are actually correct. Uh the model itself is robust.
The model is not susceptible to prom injection attack. uh the model wouldn't suddenly one day go crazy and and does the exact opposite of of what I asked the model to do. So in this case uh of course like the verification of the model itself can be done hopefully in the future uh by some uh using some clever kind of like techniques. But now how can we actually do it now? How can we kind of like have a certain level of alignment for this uh autonomous AI agents given that we can't verify the correctness of the model in its entirety.
Uh this is where I feel that by kind of like breaking the uh the AI agents into separate modules and defining clear and because now we have each module that is actually a small functional unit. It's easier for us to actually define what is the correct behavior. What is the boundaries of the behavior for all these modules? Think about it as like as a human being we have laws that define what we can do and what we can't do, right? And the thing is you might want to do certain bad things but because there's a certain law there that says that you can't do it and therefore you don't or like I don't that to me is actually good enough because the the kind of like the actual bad effect is not uh instantiated into into the real world.
So on in in terms of like level five where we are kind of like establishing the correctness of the decisions and the models I I feel that this this is a worthy goal for us to achieve but currently like what we can do really is to establish rules or boundaries on what can and cannot be done and enforce it for our models. So with that we kind of like move on to like uh proof of agents with TEES and how uh how we can use TE's to do it. I've actually touched uh on on it like earlier in my presentation. TE's provide us with confidentiality, integrity and we can make use of it to perform like the uh to provide attestation of the proofs of the environment of the model itself and we can make use of the TE to actually enforce the rules that we want the boundaries of what our model can do. And if you think about it, let's go into something more concrete like for example, how can we prove that the model is indeed correct?
We can use like in general two ways. We can prove that the model is using a specific model through uh TE proxy or we can directly just run the LM within a proxy itself. And what this allow us to do is to kind of like have this uh have this ability for us to verify or to verify that hey the request that I'm sending out indeed comes from say chat GBT or indeed comes from OpenAI or from Entropic or from any other uh LM providers that you would want to use. Next I can actually run a T uh sorry run an LM in a T itself and this again there are two ways on how we can do it like remotely uh this allows us to run 671 billion models in the cloud I mean if if you have the uh money to actually pay for it uh instead of running it on our home machine and because it runs in the TE it doesn't matter if you actually own the hardware or not you can have confidence on the confidentiality and in and the integrity of the model or you can just run it yourself locally. You can buy a GPU, put it into your desktop system and run the model locally.
And one thing that is very interesting is actually like all of us here have a TE. Our phones actually have a most of our phones. If you have an Android phone or Apple iPhone, you have a TE that's running in your phone. And we have actually tried running certain uh lightweight models uh for example uh uh 3.2 billion parameter llama model uh on the phone itself and to use it to actually perform certain task using Eliza.
If you're interested, I can share more. And kind of like with all of this, the question is what can we do? Why do we need this ver verified AI agents? One, the user have increased trust and confidence uh when they are using the AI agents and two we can actually provision services for all these AI agent. For example, we can provision uh free RPC access for AI agents.
We can provision free transaction like guestless transaction for AI agents. AI agents can can have certain u autonomy to select for certain charities maybe in the future. So I I think these are all very interesting and uh nice use cases where if we are unable to verify the AI agent, we are we are not or there's a very high level of risk if we want to use the AI agent for all these use cases. Uh I'll kind of like skip this. So really tees bring to AI agents uh kind of like two main things confidentiality integrity and by using uh and by using that we kind of like we can ensure the autonomy of the AI agent that there's no back door that it performs to a certain extent uh what I want it to perform and it shouldn't do to a certain extent what I don't wanted it to do.
Uh with that uh that comes to the end of my presentation today. Uh we have some socials and uh stuff like that. So thank you. Thank you for your [Applause] Thanks. Uh okay.
Okay. Go ahead. Um I think a little bit like Okay. Um does anyone have any questions? Nope.
Thank you. Okay. Thanks then. Um and for [Music]
Automatic transcript — names and jargon may be misspelled.