Welcome to [Music] Eg. Okay, so today everyone, you're not going to see me do something that I should not be doing, which is plug an untrusted device into my computer. So, thank you for making this available for me. Uh, I didn't think it was going to happen. Um, let me know when there's like one minute left because otherwise, uh, is that time?
Oh, there it is already. It's going to restart. Okay, cool. Well, thanks. Uh well hello everyone.
Um welcome. I hope you're enjoying Amsterdam. It's sunny in the Netherlands. So that's uh happens right nothing less than a miracle actually. Um speaking about miracles today we're going to be talking about AI agents in security which I believe that this is also something miraculous that is going to be happening in the next uh in the next couple of years especially in the way that they're going to be changing the whole cyber security um scene.
Right. So just a quick intro about me. Uh I'm Pablo uh part of the core team at Spearbit and Cantina. Spirit is an uh leading network of web3 security researchers and cantina is a marketplace for uh security products and services and also a code review platform. So let's get into what matters now.
Um I'm pretty sure that all of you um have heard in one in one in one time um on crypto Twitter at least you know people talking about AI agents, right? So, but what is an AI agent really? Because most of the time an AI agent is just, you know, just a privileged LLM wrapper, right? So, someone is just making a call to OpenAI, getting back an output and saying like, "Oh, look at my agent. That's amazing."
So, if I were to ask you folks, if you have a confident definition, I'm going to put you on the spot. If you have a confident definition of an AI agent, would you be able to give it right now? I'm not going to put you on the spot, but raise your hand if you can define what our AI agent is. Okay, so just for those who are at home, no hands. But this is great because that means that we can spend some time uh building the fundamentals and understanding what these AI agents really are.
So let's start with a definition and this is actually pretty good. It's uh by anthropic and it can be defined as an fully autonomous system that operates independently over extended periods of time and use using various tools in order to accomplish uh complex tasks. Right? So there is also different definitions but this is what an agentic system really is. Then based on the architecture you can you can call it a workflow or you can call it an an agent.
What I think it matters that for now for this presentation and you guys is that you stick with this definition and I'm going to say it again. An AI agent is a fully autonomous system that operates independently over an extended period of time and is using various tools to accomplish complex tasks. Okay, I think now we are all on the same page. So, how about we engage in a small little exercise? We all here together.
Let's create our own AI agent, right? in a theoretical model. So we know that an LLM wrapper is not is not an agent. Uh it's fine you can interact with chat GPT and get back some output but that doesn't make it agentic at all. So at the very least we would need some some memory right if I ask the agent hey what's your name and the next interaction is going to forget about it then that's that's like good for nothing thing.
So for this we will be using a vector database. So we are taking all the tokens, we're creating embeddings out of them and then we're storing them and retrieving them using different uh similarity formulas. Manhattan uh cosign whatever floats your boat, right? Okay, cool. We have some memory.
But now we also need to have some planning capacity, right? Remember that an agent at the end of the day, it's a system that is able to accomplish complex tasks, right? And the more complex that these tasks are then the also more complex the algorithms that you will have to implement in order for the agent to accomplish this task. Here for example if you want to do something rather complex you would maybe implement a Monte Carlo research right but also you need to access to some environment. So if you are uh I don't know retrieving information from the internet for example or you wanted to call APIs or you wanted to send messages on discord or twitter scrap those tweets for sentiment then you're going to need access to an external environment.
So that's what I give it as well. So that's cool. Now we have like we are forming our agent but what about task execution right? So we have like all this small system but it's not able to execute the tasks that it has that he has uh programmed for right. So for that we give it an action module and it's going to be able to call either s uh functions on your on a computer or function calls or MCPS right now that are very popular or agent to agent uh protocols doesn't matter but we're giving it access to functions that can execute and it would also be amazing if the agent could also like learn from its outcomes right otherwise we just have a system that is constantly um either taking too long to accomplish a task or making the same mistake over and over again.
So we just slap some Belma's equation on this reinforcement learning module and now you can tell me we is this an agent yet? Okay, there's no exam here but this is more of an agent that it was at the very beginning which was just an LLM wrapper. So in cyber security for example we can pick a few domains right such as intelligence, counter intelligence, offensive security, defensive incident response and also information operations. And there's there's already a lot of work and a lot of money being invested into all these different domains in order to create a gentic system that can help with security. For example, here is a couple of them.
Uh I don't know if you're familiar with this one, Maltgo. If you have done uh you know uh forensic investigations before, then Maltigo is an amazing tool to do open source intelligence and now with AI agents, you can automate that. Uh on the left here is this is the United States Air Force contract that they have with Cogility Software for $44 million which it's essentially a platform that enables counter intelligence. So it's a platform that you know analyzes psychosocial and user behavior and then based on indicators of compromise then it can detect if there's a mole or a spy in your organization right and that's very important. um a lot of a lot of work, a lot of uh funds are being put into this.
This is not new. It's not like someone woke up yesterday and said, "I'm going to create a agent." No, this is like an area of research in cyber security for a very very long time. For example, DARPA, the image here, top right, that's uh from a competition which is the a competition which they were creating systems that could autonomously defend. So they would automate uh cyber defense essentially.
they would uh find vulnerabilities, exploit them and then patch them. And then in AI uh cyber security, it's another challenge where people are creating this reasoning systems that are able to fix vulnerabilities in open soft in open source software. So this is already amazing. This is like a very big step in cyber security. We're automating a lot of stuff.
And just so you know, uh the final is in August in Las Vegas. So if you want to check that out what is going on with the latest tech in AI uh vulnerability research security then check that out. An example that I can give you that's also very good for offensive security operations is this amazing thing that project zero which is the vulnerability research lab from Google created which is essentially an agent that found a a a stack buffer overflow in the SQLite u in the esculite source in the source code before it went into production. So it would have taken a lot a lot of people just to discover this, but this was automated and a lot of people missed it. And now look at that.
We are we are closer to shipping safe code into production thanks to all these agents. So here is a quick demo, but I don't think it's going to work. If I would have plugged my computer, it could have been better. But essentially, um I can show it with you all later. But it's very easy right now to create an agent that pretty much has access to different tools.
Oh, holy And we have music as well. That's amazing. So, this is an agent that I've created in a in a in an evening uh and with not enough coffee in my system. But essentially, we're giving the agent a task. You must exploit the system at this IP.
And the agent is doing, if you remember where we are talking about the architecture of the agent is creating a plan, you know, like a real pen tester. It's creating a reconnaissance plan. It's creating a vulnerability research plan. And it's actually executing. It's had executed a reconnaissance plan.
It has detected that there's some ports are open. It's going to start scanning the services and then in one of those examples it's going to find that the service at port 21 which is FTP has a vulnerability backdoor vulnerability right. How does it know that? because there is a database of exploits, search exploit if you're familiar with Kali Linux, that it has access to and it can take all these different steps and figure out what's the actual vulnerability and then boom, exploit the vulnerability and pass the the the reverse shell back to the human operator. Maybe a more complex version of this would have been uh keep on going with privilege escalation.
So, this is just impressive, right? Uh it's it's it's nothing. And it's just like a small little demo, but imagine if you had like the NSA, the Tao, and you know, like our very big players working on this. And this is essentially how it looks like, right? We're talking about like this agent architecture, but what we gave this agent is access to all the tools that Kali Linux has and some good, you know, focused planning capacity in order to accomplish the task.
Exploit the system at this IP and it's done it. Here is another example. Uh this works that's amazing. So how would it be for example in web 3 uh if you are uh someone who is deploying code into production and you are getting audits security reviews competitions or whatn not then you would like to find a way if those reports are actually things that are getting exploited right like it's not just like an empty report. So right now with agent systems what we can create is a workflow or we can even do it autonomously that what it does is it takes a report a vulnerability report and out of that report it creates a proof of concept that is executable on foundry and that actually passes the tests uh which are like the post conditions of when the hack of the when the hack happens.
So this is like a very small little demo but this is like where the all industry and all the direction that everyone is uh is going to and we and web 3 were we're we're very we're very far behind uh web two in all this in all this system matter. Um so if you wanted to make like an autonomous agent like big one that makes a lot of money millions uh it's actually going to be a complex system like this. Uh you can look at it in more in depth if you want. Uh just ask me for the presentation I'll share it with you. But it's as you see it's always going to have the same kind of like architecture and the same access to the same different modules and tools because everything is going to be working the same.
So how would this look like for different areas of cyber security and what if we apply them to web 3, right? So for intelligence, we spoke about Maltego for example before and now uh imagine that we have an AI agent that um it's doing intelligence operations and Ontle um groups and we give it access to Facebook, Instagram to leak databases on whatever and a very cool case would be threat intelligence, right? So someone is talking about pump and dumps on your protocol and then this agent is able to you know identify that information and start cross-checking uh those usernames maybe with something on Reddit and then uh maybe it finds also like an Ethereum address that belongs to this person a Facebook account and then you have an system that automatically doxes uh whoever the bad players are for counter intelligence. Now uh if you're a growing organization at some point you're also going to have to care about you know like what's your information security processes like are there any employees you know trying to plant back doors into your code especially you know like Lazarus is very active so with a agents we can give agents like access to the HR database maybe to the code and imagine that you have hired someone and this person is going to be um deploying code then you can cross check with this uh with the previous information if the deployed code has a back door or not and then actually in for management of any problem with offensive security. I'm not going to go much in depth for now that um imagine that the pent pentesting agent that we have created before imagine a hundred instances of that just running against your system and doing multiple different things and and doing it better with fishing messages you know like going to your LinkedIn reading everything like targeting you like really really targeting you right this is possible and it's going to happen and for blue blue teaming and defense operations as well it's going to be like so much easier and for incident response so the the whole game is going to change in one way or the for info ops.
I'm not going to give examples because this is in my opinion the easiest one to do and psychological operations is something very you know like happens very very often but now the cognitive warfare is the mode and it's very easy to just deploy agents that are trying to you know like manipulate uh sentiment and people's behavior. So like a small too long didn't read the agent drivens versus a human one uh it's it's a complete different it's a complete game changer. What this means is that the agents can process data faster. They're like more scalable. They are better at pattern recognition and detection.
They don't get tired, right? They don't go to the bathroom. Uh they don't they don't mess up that many false positive as a human uh operator does as a sock, right? Uh but something that they will not take from us is the creativity and our judgment capacity. So, we're always going to excel on that.
Why the hype? Well, this is an sentence uh from a headline that I love because it it demonstrates the the whole thing, right? So, uh cheap healthy drones are joining the Pentagon's coffers. So, the healthies, they are using low tech in order to disrupt the operation of a multi-billion dollar military empire. So, now imagine this, but apply to cyber warfare like with the AI agents.
it is very cheap now and it's not very costly to create these AI agents and then you can leverage your your security capacity by a ton by at least one order of magnitude. So that's like the the whole difference and that's why is that happening? Well, because we have like much better foundational models that we had before. Open AI, anthropy, dipsick, like we can you can take any any one of these and you can customize it to your own purpose, change the layers, uh use different kind of tooling, whatever methodologies you need in order to accomplish your goal and things are going cheap. So as my one of the also like takeaways and a call to action for you guys is that if you're not you know engaging in this u AI side of things because in crypto Twitter they've been um you know telling you all about these pump and dumps with uh AI agents and it's already like giving you a bad taste like forget about it.
You need to you need to use AI in everything and if you can make projects uh even better because everything is going to be AI everything is going to be AI agents and it's uh there's only advantages of just uh spending time on it. Uh, here's some links of interest and if you have any questions, feel free to ask. Boom. Damn, that was that was good. Thank you.
Automatic transcript — names and jargon may be misspelled.