Does Static Analysis Actually Help AI? — Michal Převrátil | Ack3
ETH Belgrade Community·Tue, Oct 6, 2026, 12:00 AM
Transcript
Great afternoon. So together we will try to figure out if static analysis actually helps the AI or not. But before we begin, I want to make sure everyone's on the same page. So let me give you a few examples of what I mean by static analysis. I have a few examples there.
The first one is find references. Maybe the first one you can think of. we have some kind of variable for example owner and we are asking ourselves where the variable is used or maybe the AI wants needs to know where the variable is used and maybe it's only concerned about the the the mutations the rights into the variable or maybe just reads anyway this is something that can be easily provided by static analysis tools we can just find the references you can filter by the usage of that reference and yeah this is just very basic basic thing that can be provided by static analysis. On the other hand, if we didn't have the static analysis, just pure AI, the AI would likely use something like RIP grab, which is a common line tool for finding some text segments across files. And in the end, at the end of the day, it will also find some references likely all references and very likely even some references that actually don't refer to the variable we are searching for just because it will match any text that is present in the workspace.
Another example, uh get colors. Uh let's say we have we have a function called withdraw. And that function we are searching basically for all the locations where the function is being called. For that we can use a call graph in static analysis. And again a pretty basic stuff.
While what will the AI actually do is again use rip grab to search for the usages of that of that symbol actually with the opening bracket also to just match the the calls to the function. And well this way it will also find likely all the references all the calls but it very much depends on the language on the context and maybe some edge cases. For example we can store the function into a function pointer and execute the function pointer later and in that case the AI wouldn't find that specific call just because of weird syntax. So maybe an edge case but there are already small differences in between static analysis which is usually very precise and AI which just takes the approximation and just some basic tools on the com on the command line in child that are available on any machine. Uh another example is to get state changes.
For example, in contracts, we may be curious or the AI may need uh all to list all the state changes that are happening inside the function and also the functions that are being called from that function. So like recursively across all the all the functions being called and for that the AI will basically need to list all the source code. So it will need to print all the lines of the function. Also, it will need to follow all the function calls into the nested functions and it will somehow need to evaluate what those lines mean and what they actually produce and if they make any state changes in the contract or not. Uh one more example if some function is reachable or some line of code is reachable from another location.
Uh again this is quite simple in static analysis and quite precise while the AI well it needs to search for those symbols and somehow it needs to read the relevant parts of the text of the source code and evaluate if it's actually reach re reachable or not. But this is again just based on approximation or just maybe let's say on the luck how well the AI can read the source code and find the right lines in the source code or not. Uh one more example get storage layout. So yeah again I think it's pretty simple and again it will just grab it will search for some text segments. So at the end of the day, most of those static analysis tools can be somehow replaced by just searching in the in the workspace using grap or rig grab and obviously we need to read the text.
We need to read the the source code and for that the AI will use set or similar tool just to print the contents of of the of the files. But well usually when we are working with the source code the AI already reads or has read the source code just because it works with that. So doesn't matter if we doing static anal if we if we doing analysis of the code of or if we are developing the code either way the AI likely already has the source code in the context. So the question is does the MCP/static analysis whenever I'm talking about AC MCP which is a protocol a standard for providing uh tools not only static analysis tools but tools in general. So whenever I'm talking about MCP in the context of this presentation I mean static analysis MCP.
So does the MCP/static analysis add any extra if we already have the source code in the context which we most likely have or let's say the major part of of the needed source code is already in the context. So does it add any extra? Well, it definitely adds the cost. So we did an experiment of 168 prompts and we've tried those prompts with two diff uh three different models GPT 5.6 Luna, Terra, and Saul.
And we've evaluated the cost without this data analysis MCP and with this data analysis MCP. I haven't described the exact data set yet as well as the MCP server. I will explain later through the presentation. For now, you will have to trust me. But as you can see, there were some extra extra dollars, extra costs for the MCP.
With Luna, it was around 14%. With Terra it was around just 1% which is very interesting. And for Soul it was around 13. Oh you cannot see you cannot see the slides. All right.
So now hope you can see that. And yeah it was 14% for Luna, 1% for Terra and around 13% for Soul. All right. So this is quite interesting. Why is that cost uh increase that small only for Terra and why it's so huge for the remaining two models?
I think this should explain some of those questions. So uh as I said before we had 168 prompts questions with with those models and we've asked our ourselves in how many times did the the AI use the MCP server at least once. So out of out of those 168 how many times did the AI used MCP at least once? And as you can see for Luna it used it almost every time. For Terra it used it only like in one/ird and for soul it was more but still still less than Luna.
And I think this explains the cost difference especially for Terra because for some reason and this is quite interesting honestly I don't know why but Terra didn't want to use the the MCP even though it had the same conditions. It had the same context everything but Terra just didn't like to use the MCP the state analysis tools. So that's that's the reason why the cost difference is so huge for Terra. Now you may be wondering why there are extra costs. I will try to explain.
So let's say that we have some kind of prompt. Along with the prompt we have the system prompt. We also have the definitions of the tools. Now it doesn't matter if it's MCP or some kind of native tools like a shell or something else or to-do list. It basically can be anything.
And we've sent the prompt. Then the agent did some kind of reasoning. And after that reasoning, it has decided to perform a tool invocation. In this case, it's very general. So in this case, it doesn't need to be static analysis tool.
It can be any tool. It can be even a shell. It can be anything. But the agent has decided to invoke any tool. Obviously, we've taken the input and we've executed the tool.
It may have provided something. It may have failed. Whatever. Either way, we've collected some kind of result. And to follow up, we need to return the result back to the LLM to the AI.
So what we have to do is we have to do we have to take the whole prefix which means the whole previous prompt including the reasoning including the tool definitions including the system prompt basically everything we've built so far in the conversation and we need to send it again to to the provider of of the AI and very likely this will hit a cache and so those tokens will be built as cached input tokens that don't cost that much but still they cost something. It's a small value but still still something and then let's say that we've sent that and the AI after some time has decided to do one more call and it will do so and when it does again we execute the tool and we need to return the result and with the result again the common prefix grows and we have to send it again to the server. So as the conversation continues and as we make more and more calls, we we well we keep hitting the cache likely, but we still get build for those cached input tokens. And if we make a lot of calls, the bill can be quite huge. And this is the reason on the previous slide why there was extra 14 maybe 13% just because there was a lot of tools and we kept sending the same prefix or even longer prefix in the over the over the time.
So how we can reduce the cost? One possibility, one option is that we can ask the AI the LLM to batch multiple tool multiple calls or two tool calls into one batch. So if those tools tool calls are independent, we can ask the AI to batch them into just one big request. Let's say we can then execute all of those tool calls at once in parallel or in Siri, doesn't matter. But after those all the calls we will then return the results back to the server and because of that we don't get built for like let's say three uh three different uh common prefixes but we will get built only once.
So this is one possibility how to how to lower the cost. And another one and this is not from my head I have to refer to Tibbo on X. This is one of the most famous profiles now. He's from OpenAI. I believe you know him.
And he says or like in his response to why they have some number of of context window limit in in codeex he replied with with this graph. I will try to explain. So basically what he's saying is that we obviously can specify the context size limit which is basically how long conversation without some summarization we can we can keep in the the in the let's say session or the chat and after we hit that limit we have to summarize we have to compile the conversation take only important points and not start over again but start with summarized summarized context and what's interesting thing is if we lower the context size window there's as I said the the tool calls and the prefix that keeps sending over and over again and since the context size is lower the common prefix also will be shorter in average because we keep building shorter common prefixes before we send the tool call and so the cost in total with respect to tool calls will be lower and in this example I believe there's 200k and 300k uh context size window there was 40% difference in in the cost which is quite a lot. So this is another approach we can take. The question is if the context comp uh compaction will will affect the quality of the response of the AI or not.
This is another question but uh from the perspective or the cost this is something we can definitely try. One more small tip. Uh most or a lot of MCB servers I can see in the wild. Uh basically provide the response to the AI formatted in JSON. So this is something you can see on the left side.
So they just format the text in the JSON which is well okay. We can clearly see what those fields mean in practice like what's what's the the description of each field. But the AI is not machine in the traditional sense and so it does not need to accept uh like machine formatted outputs or inputs let's say and we don't need it to provide JSON. So what we did in vague which is the tool built inhouse in EC3 is originally we also had the outputs in JSON but then we switched to custom format basically plain text with custom formatting and as long as the output from from the tool is clear and it's easy to understand what what each of the text field means in the in the output. It's better to to use the custom formatting just plain text because per our evaluations the number of tokens has reduced by 58% on average just by using a custom format.
So this is also quite interesting and it makes sense because we don't keep repeating the the names of the fields and if it's obvious what those fields mean why we should keep repeating and using JSON just because it's a practice in like traditional traditional software. All right. So, we've talked about costs, but is there any measurable in impact except for the cost? In my opinion and based on my experience, it's compiler bugs. Uh, in a happy world, there would be no compiler bugs.
But, unfortunately, in the context of solidity and soul C, there are some compiler bugs. For example, this one quite infamous one, the call data encoding bug. I don't want to dive into details but basically under some strict conditions and in some versions of the solidity compiler the code will will behave completely differently than how it's supposed to and how it's documented and this is something that that's unfair with respect to auditors but also with respect to the AI because the AI also reads the code and if it's not aware of that compiler bug it will just miss it and I'm not saying that this specific bug won't be caught by the most recent LLM models just because it's so famous that maybe in some prompts and in some cases it will be still detected. It depends on the luck and on the size of the project and so on. But all I'm saying that it's unfair and the LMS may miss it as well as auditors if they are not aware.
And what's even worse that there's 59 solidity compiler bugs in total and so it's quite a huge number. Obviously, we could try to guide the AI to focus on each of those, but the cost would be like growing with each of the compiler bugs. And so maybe it's just easier to use static analysis to detect those compiler issues because usually it's quite easy to write a detector for those bugs and we can save a lot of money and have higher confidence with respect to to these compiler compiler issues. All right. So is there anything else measurable except for compiler bugs and we are getting to the b main benchmark.
I will try my best to explain and so we took one of audits we did in past. It was a proof of stake client written in go. It was quite a huge project and we've extracted 138 questions and those questions were basically very specific questions on some attack vectors in the codebase in what what I mean by very specific it was like in the file ago there is a line something and we are assigning assigning some value to some variable cannot it lead to a crash because of something so it was very specific and the answer was expected to be yes or No, it is a bug or it is not a bug. Maybe some extra comment that's not important for the benchmark. The the answer could be binary.
And what was the split of those uh uh questions? uh I believe that 55 of the questions were were uh expected to be marked as valid which means that the the the issue was truly marked as valid by humans while the remaining 79 79 questions were expected to be invalid which means that human uh flag that this question shouldn't lead to a valid issue in the data set. So this is the data set and then we've we took three different agents you already know GPT 5.6 soul terra and luna and we've launched those exactly same questions in the same environment on the same project with two configurations with the MCP on and off and then we've compared the results. Uh then to explain the MCP server what exactly it had or it didn't have.
So it had some basic tools to refresh the index because the AI may have changed the code especially if it buil if it tried to build built some kind of proof of concept scripts. Then there was some basic tooling for listing the symbols the the packages and also to list the diagnostics from the llinter from the built-in llinter in the go tool chain. uh there was uh tool for uh finding references and listing symbols in a file and there were also tools for uh finding uh the the colors and colleies in the call graph. So let's check the data. There will be a lot of slides with the data.
So we will go through each of them. So first of all we've measured uh out of those five 55 questions how many were confirmed by by the LLM which means uh out of those 55 questions how many were answered correctly and there are the three different models and there's the percentage like on the x axis so on your right there's higher score towards 100% % and there are three models and there are two dots the blue one which is with the MCP on and the black one with MCP off and as you can see uh the slide says that there was no improvement in sensitivity which is in the end true but there's still some difference especially for the first two models and the difference is actually negative so with the MCP the outcome was worse while for Luna it remained the same this is quite interesting ing in my opinion and I would say that most likely this was just statistical noise. So it's not a significant statistical result. I will talk about that later but for now we can see that it made no difference for Luna but there was a small difference for the two uh other models. Then we check out the precision.
So uh yeah now we can see that the the MCP actually made difference in positive sense. What this slide captures is out of those issues that were flagged by the AI as valid. How many of them were actually valid? So this is quite a different metric because the previous one focused on the known set of 55 questions while now we have those all questions marked as valid and we are comparing the comparing basically uh comparing how many should not have been flagged as valid by by the AI. And here we can see that there was almost no difference for soul while there was quite significant difference for Terra and for Luna.
This is just a combination of the two previous metrics. So I think this is not super important but basically you can see that it made almost no difference for Soul and for Terra on average when we combine those two metrics together. while there was positive impact for Luna. And the left part of this slide you've already seen before. So this this is just the the repeated data from the previous slides.
While on the right part we describe uh what's the refusal rate or how the MCP helps AI to refuse false issues or like f questions that lead to false issues, invalid issues. And there we can see that again there was almost no impact for soul but quite significant impact for Terra and Luna. So it seems that the MCP helps uh those two models uh to Terra and and Luna to not over report which means not submit that many issues that in the end will be false positives. So this is something that's interesting and let's take a deeper look at that and in this case we are focusing only on Luna because this was like the the the biggest impact from all the data because we've seen that uh MCP didn't help Luna to to report more to positives more interesting issues that were valid but it helped not to over report not to report too many false positives and there's the with MCP on and off where you can see that out of those 79 questions that in the ended into issues that were not valid there was four questions difference between MCP on and off and if we use some kind of static analysis uh stat statistics instrument called 95% interval and something and we just apply it we can see that the outcome is still not statistically statistically significant which means that we can say maybe MCP has helped especially Luna not to over report not to report too many issues but it's not statistically significant but maybe there's some kind of influence. So these were very specific questions.
So we we thought what if the problem was in specific questions let's try something else. So we will supply open-ended questions to the models and we will try on that scenario or on those kind of questions. So we've taken the same project the same issues. The slide says 45 not 55. The reason is that in those 55 questions there were some overlaps which means that some of the questions led to the same issue but in the end the number of issues in the project was 45 and what we did is that we've launched very simple prompt basically perform deep security analysis with those three different models and we wanted to see if the MCP will make any difference on discovering issues those already known well-known issues in the project or not.
And as you can see again the difference was not so significant. All right. So we said maybe it doesn't help to discover new issues that much. But what we what if we tried uh uh to use the MCP to build something because that's quite a different task uh review versus development. So we've taken those a subset of those known issues and we've gave them to the AI and we've asked the AI all the free models to build a proof of concept script for those already known valid issues and this is the result and again as you can see there was no significant result and in the end the MCP didn't help there was actually one less proof of concept script generated compiled and basically passed which is again very likely statistical noise and not something that's super statistically significant.
So we may be asking ourselves what if we changed the definition of tools in the MCP what if we give it gave it a completely different set set of questions and different tasks to focus on and this is basically open-ended as open questions after this presentation because it very much depends on your MCP on your statalysis tools also on your workflow and what you are trying to achieve with the AI but TLDDR of the of this presentation. Do not assume that the MCP which is basically or static analysis which is more precise will improve anything in the final outcome of the AI or that it will reduce the cost because when we just take a look at those basic MCP tools for navigation across the project the numbers say that maybe it won't improve anything and it will just introduce extra costs. So yeah, that's all from me and uh make sure to follow me on on X on on Telegram if you want to know more or just you know uh talk about that deeper. The tag is Mitch Prep and also be sure to follow ECFree AI on on X and yeah that's that's all.
Automatic transcript — names and jargon may be misspelled.