# Integrating Blockchain Data into AI Agents

- Channel: [ETHBerlin](https://streameth.org/ethberlin)
- Date: 2025-06-19
- Duration: 25:27
- Watch: https://streameth.org/watch/68553c5590bd41297b4fd7f9

## Description

AI agents are transforming how we interact with data, but their effectiveness depends on real-time, reliable information. In this workshop, we will explore how to integrate blockchain data into AI agents using a SubQuery Indexer SDK. Participants will learn how to query on-chain data from different networks, process it efficiently, and use it to enhance AI-driven decision-making.

## Transcript

And the motivation for this workshop is essentially current complex and massive data that we have to deal with if you want to be working with historical data and make the analysis of it. And it's really hard to query directly and even if it's possible, then you still have to have a certain understanding of data structures as well as the, yeah, you still have to do some structuring. And you also have to use some tools, but like you can also leverage AI as a tool because it can help interact with such structured data with a simple natural language request. So that as someone who is not technically a business owner who wants to get the insights from ancient data, you don't have to worry about knowing those technical intricacies. So by combining blockchain data with AI, you can get like user-friendly querying tools. So the goal of this talk would be to demonstrate a practical integration of using open source tools connected together in a completely local setup and provide an accessible starting point for exploring data exploration with AI. A preamble, this is going to be a beginner-friendly overview prepared by someone who is not specializing in AI. So if you know more, if you can suggest some of the things for further research, please do after this workshop. I'm a full-time student and I want to be the same as one. So there's going to be no, there was no custom development involved during the presentation of this workshop. It's just those open source components stuck together and working to essentially make this proof-of-concept. And obviously, it's not production-ready, so you shouldn't be going to that. You shouldn't be pushing it to your production app. So this is what I'm going to showcase. So this is a process of the things that I'm just going to be showcasing. So we're having a, in parallel, we're going to be running a local node. We're going to be indexing certain data from it, and in parallel, once it's going to be indexed. So indexing essentially, you know, it's structuring, the unstructured, blockchain indexing is structuring the unstructured data and saving it in a certain structured format like a relational database. So on the right-hand side, we're going to be doing that. And in parallel, we're going to be creating and using an agent. I'm having it in double quotes because it's not going to be really an agent because the agent is able and capable of making the decisions autonomously, but like in my case, it's just going to be an AI-powered tool that's going to be just translating the natural language to SQL and then interpreting the output from the database to a natural language result. So it's like easily readable. And this is what I'm showcasing on the left-hand side, so here. So essentially, this agent is going to be creating a SQL for further querying. Then we're going to be manually controlling it, whether it should execute the SQL command. And finally, it's just going to send the results to the EVM, and then EVM explains what exactly the outputs of that. So components that we're going to be using, we're going to be running a substrate node with valid revise and revise-lpc-configure. This is a RISC-V virtual machine. I'm going to be spamming the smart contract with a lot of data, and I know the RISC-V is going to be capable of handling such kind of data, so I chose that. So the PDM is not a part of this talk, so be sure to check the talks of the fellow speakers on that part. I'm going to be spamming with hard hat, going to be using substrate indexing as SQL for indexing. I'm going to be hosting LLM on Allama, and in order to build this agent, we'll need an OpenAPI-compatible endpoint, and Allama gives this to us. And as a framework for building an agent, we're going to be using BlankChain itself. So, demo time. So, we're going to be starting that with hard hat, let's start a node first. So, I'm building a substrate node with valid revise, so this is the contract, we also need a proxy server so that we can deploy the contract and interact with it. So, for that, I built RPC in advance, and it's a very simple contract that we're going to be dealing with today, it's like that, so essentially it's just a number, an event number stored, and actually a single function that we're going to be using, store, so it's going to be accepting a number, save me into the storage, and then a certain number stored event is going to be emitted, so that's it, very simple. So, we can go ahead and run the node. So, I have started it, we can also check if it's running, in case you don't know, this is a Google CloudJS UI, could be usually considered as a blockchain explorer for substrate and blockchains, but I'm having issues with the UI for some reason, I'm not seeing the block, oh yeah, here it is, so the node is up and running, so we can start spamming it. So, I use a hard path for that, I'm going to be deploying the smart contract that I've just mentioned, and then we're going to be just interacting with it, so this is the script for interaction, we're having an infinite loop here, and dealing with infinite loops, and yeah, just going to be taking a random value and call the store function with the value. So, this is command to do it, so I'm delaying my previous deployments first, I'm deploying a contract, and then I'm starting the script. So, if I start running it, having the contract built, and yeah, we should have it started anytime soon. In the meantime, I can proceed to the next part, probably, oh yeah, that's already started, so we can see that we're started spamming, and in the explorer we're also seeing the events appearing, so there's a lot of data that we can actually index and work with, so it's cool. What's next? Once we have nodes started now, as well as the spamming of the contract, we can start setting up the indexer itself. So again, an indexer is a service which is just structuring unstructured blockchain data into something that could be worked with. So, this is a folder for the indexer, so that's going to be a pretty simple setup. So, this is a file to essentially specify a data schema that we're going to be dealing with, so since it's only a number stored, only a simple number that we're going to be storing, that's going to be a single entity. So, as you can see, it's block height, number, and sender. This is the only data attributes that we're storing for each transaction. A second thing would be to configure the actual endpoint that we're going to be working with, so it's locally hosted node, so that's the endpoint, and the chain ID is that. You can trust me on that one. And what's going to be done is we just want to handle this specific event, which we've just created in a smart contract, and pass this event as a parameter to the handle lock function. So, this function is just a writer to the database, essentially. So, as you can see here, it's creating a new entry for, once it's seeing a new event, it's just creating a new entry in the database. By ID, it's using the transaction hash, and for number, it's using the value being used. So, if we just start it, we'll see that and hope that the index is going to be doing fine. So, yep. So, what's happening behind the scenes? It's really simple. We're just having a new database with a specific, a single table called number store. So, what we just did is essentially structured blockchain data to something, to a relational database. That's it. And since we have a relational database that we can work with, and since this is, sorry for web 2 interfaces here, actually. That's probably not something that you would expect from the pre-conferences, but yeah, it is what it is. And since, yeah, it is a, yet a popular thing that a lot of companies are dealing with, there are, of course, a lot of out-of-the-box tools that could interact with that, including an agent SDK that we're going to be using now. So, essentially, what we want is queries as data from the LLM, which is a simple thing because a lot of agents are capable of interacting with that. So, proceeding to the agent side of things. So, of course, all of that is locally hosted and done from my machine. It's not the most performant machine ever, so it's capable for doing this local analysis, which might be useful for someone who's just, who don't want to be dealing with a lot of computations. So, what's next? We start the agent itself. So, let me go over the code of the agent. Hopefully, you're seeing it. So, I'm just going to be briefly describing what we're going to be doing. So, we are starting an LLM. I'm using OLAMA, and I've picked this random model from the library. We're going to be connecting to a podcast database that we're hosting in Docker for convenience reasons. We're also providing certain instructions in the form of a prompt together with that. We're passing a schema of the database so that LLM knows how to design the SQL query itself, and we're also passing a question, of course. So, here are some of the functions that we'll need, but essentially, it all comes through three steps of this agent. This chain of commands is to write a query, execute a query, and generate an answer based on the initial question. So, those are the steps. They're going to be executed sequentially. So, yeah, we can start running the script, and I'm also I'm also having a few example queries prepared here. So, the first one is going to be how many times the number was stored so far. I'm also giving it a hint to use for SQL, but let's try without that. So, what's happening? It's trying to generate a SQL for our questions here. So, we can visually see and validate it. So, it seems the app metadata, join app metadata, for some reason wants to join the app metadata. Let's try to execute and see if it works. It's zero. Wow. All right, let's try again. Probably, we don't need partitions by. All right, let me try to return the hint. Maybe it will help this. All right, yeah. So, with a little help, it worked. Again, it totally depends on the LLM you choose. You have to probably teach it by providing an inference, but like for a default one that I For instance, that one, what are the most frequently stored numbers and how many times was each of them stored? This is the SQL we received. Should be fine to me. Oops. Yep. Failed to execute. Cool. Let's try again. Oops. Don't use a full selection. Come on. So, the reason it's failing is because mass query ID. I don't know why it's doing that. All right, last attempt. Oh, yeah. It worked from the start. So, yeah. Here is the numbers and here is the interpretation. The most frequently very convenient for, again, those who are not really tech savvy and cannot even work with or don't even have the time or possibility to work with any tools or dashboards. Yeah, there is a tendency to switching to chat-based analysis. Let's also work with the third query just to test it out. What is the last stored number that got stored? I'm also hinting it again. So, use a block height. So, yeah, it worked from the first time. So, it's cool. All right. So, and on that point, I think this is all that I wanted to show you. So, if we have time for, I don't know, open conversation discussions, any suggestions for further research or any things that I should look into as someone who is just yet experimenting with that, I would have to hear your suggestions or maybe answer some questions. Thank you. It took less than an hour to configure all that. Well, you have to understand like probably the hardest part to configure is the indexer because you have to know what data to deal with and how to handle it. But apart from that, it's easy. You can use it out of the box. Yeah. The context, you mean. Yeah, you can set it up so it learns. So, yeah, you can set up in the form of a thread. So, it's having a context to work with just like any LLM. You can do it, yes. But do you want to use it? I wouldn't do it, no. I think it's important to understand that, yeah, well, that's your probably risk appetite that you're probably want or not. Or it's going to be a very good AI that you want to be dealing with if you want to be executing transactions. But like I wouldn't personally do it, but I'm not a person who is having a huge risk appetite. No, I didn't publish it. These are several repositories. This is the data lock, like the chain lock. So, well, I think it's, what's the target audience here? So, what we're dealing with is a chain document. It's not something technical and deep. It's not like traces. It's not like anything consensus related. It's really those like small, it's called business data points that are stored on chain, like blocks, transaction, and events. So, the target audience is like whether who is interested in only that kind of data. If you are more into, if you're in the development of consensus and you need like good debugging systems together with detailed locks, then yeah, it would be possible to set up a data lock, but again, that's just a different target audience. But the question is whether, as a consensus developer and someone who need to work with locks, do we even need AI for analysis? Yeah, then that's going to be a good use case then. So, again, I'm using LankChain here. So, if you can check and see that, I'm probably sure that LankChain is really having the data docs supported out of the box. So, if it does, then it's probably going to be possible. Yeah. Sorry? Because it's a good and flexible framework that having out of the box as a relational database adapter. So, LankChain is not the most sophisticated one, not the most modern one, but it's just a thing that is able to easily showcase those sequential operations that are needed for this specific use case. There are actually like, of course, there is a next version of LankChain, which is called LangRaph. So, if I'm having my sub-sequential, LangRaph would be able to autonomously decide like which step should I execute next in order to achieve something. So, this would be more autonomous, but since I want to simplify my use case a little bit, I use LankChain. Yeah, just for, again, for simplification reasons. Honestly, I picked a random one from the Lama library. No intentions to use. So, I'm also curious in knowing if any of you have been having any experience previously, whether it's actually a good approach to be structuring the blockchain data to anything more structured like a relational database, or is it something that we can use separately? The reason why relational databases are mostly a de facto standard for the indexes is because, first of all, it's easy to wrap the GraphQL service on top of it so that you're having those interconnected entities, and secondly, like indexers is something that is using a feature of relational database, the indexes themselves, so for faster querying, but like I was curious if anyone in the audience have any experience with non-relational databases and indexing data to storing it in non-relational databases with work with it. So, what are the So, have you been comparing it to relational database performance? Hmm. All right, understood. Thank you. Any questions from anyone else? All right.
