# Fixing Stack Too Deep | Moritz Hoffmann (Berlin Ethereum Day, June 2026)

- Channel: [Berlin Ethereum Meetup](https://streameth.org/berlin-ethereum-meetup)
- Date: 2026-09-09
- Duration: 25:49
- Watch: https://streameth.org/watch/yt-wbST1xQ-9Jw
- YouTube: https://www.youtube.com/watch?v=wbST1xQ-9Jw

## Description

In this talk, Moritz Hoffmann (Solidity & Argot Collective) explored the status quo of viaIR, dived into aspects of the new backend's inner workings, and peeked at what is planned. They addressed two long-standing pain points with an overhaul of the viaIR backend: going from Yul to an SSA-CFG representation allows us to systematically deal with stack-too-deep as well as bring down compile times.

The Berlin Ethereum Day was a one-day event held on June 15, 2026 during the Berlin Blockchain Week. The full-day program brought together speakers from the Ethereum Foundation and the broader FOSS, privacy, and security ecosystems to explore the future of Ethereum and self-sovereign technologies - from technical direction and core values to the challenges and opportunities ahead.

Future Meetups and Events: https://www.meetup.com/berlin-ethereum-meetup/

More information on the speakers and the agenda: https://berlinethereumday.com/

## Transcript

So we already heard that stack too deep is one of the most pressing issues. But I would say another pressing issue that I put in the subtitle here is the U optimizer performance or generally the performance of SOC. Whenever you compile a bigger project, you might have realized that it can take tens of minutes if you're unlucky to actually get the bite code out, which obviously is a long time. And to um increase throughput of developers, but also you know get get rid of this annoying problem. uh we are investing in a new back end which is based on SSA cfgs and yeah so let's see what that is all about uh first I want to give a small motivating example here we can see um some some u code so this isn't this isn't straight solidity code but the code that falls out of it if you go through its internal representation and that internal representation is relatively close to the EVM. So what it does it and end stores a memory guard address at 0x40. It has a function which has many return values. I put an ellipses here but it's what's that 18 return values. Then we simply take these 18 return values and store them to storage. Now if we try to compile this through source C present day we will get one of these pesky errors which is called ST to deep and that's kind of the the highle view of what you can see as a user but in order to determine what we can do to get rid of it in the end we will have to zoom in a little bit and go deeper into the internals of the compiler. So what happens if we do d- yir or if we go [snorts] through u we have at the left hand side solidity code which some of you may be familiar with and then the solidity code is through a translation layer transformed into u. So that is this intermediate language between EVM assembly and solidity. And then potentially in the middle we have a U optimizer which will at that level of abstraction perform some optimizations which will give you gas benefits and such and then we end up at EVM assembly. The stack to deep error itself occurs at the step from the UL optimizer to EVM assembly or from U code itself to EVM assembly. Uh the name itself already hints a little bit at what it is. It is a problem with the stack. So the EVM as a stack machine and UL itself doesn't really know anything about the stack yet. The transform to EVM assembly though will then [clears throat] uh have to deal with the limitations and benefits that the EVM stack brings. Okay. So let's look at that a little bit. And in particular what I want to look at is a step that we call stack layout generation or sometimes it's also called stack layout scheduling. And uh that is a step that will for each function and for each built-in call where a built-in could for instance be here an M load on the right hand side. Um we will prescribe a specific layout of the stack that we expect the EVM state to have at that point in time when executing the the bite code. So [snorts] if we if we walk through this example real quick, we have a function f which takes two arguments v 0 v1 and it returns v4. That means initially we can assume the stack top and if I look at these representations of stack it's important to keep in mind that the rightmost bit of it is the top. So v_ub_1 is below v 0. We have a stack V1 V 0 with V 0 at the top. And then we want to execute the first instruction which is an M load of V 0. And since V 0 isn't used anymore downstream, we can simply emit the M M lo code and the V 0 is going to be replaced by a V2 symbolically. [snorts] Now afterwards we want to em load from a static address 42 right this is just a madeup example. So we will have on the left hand side the preparation of that operation that we want to execute which is going to be a push of the constant 42. Then the stack state is going to be v1 at the very bottom. Then we have the v2 which is the thing that was loaded previously and [snorts] the 42 that we just pushed. Then we can execute the second M load. Then uh the 42 is going to be replaced by a V3 symbolically. We swap them around for a add that we execute afterwards. And then in the end we want to return V4. And after the um after the add instruction the state is going to be v1 v4 which is one v1 to many. So we will have to swap one pop and then we can return from the function. That's kind of what uh stack layout scheduling looks like. So for each instruction you prescribe a layout of what that instruction needs needs in terms of shape of stack symbolically. Now [clears throat] not every function is just a single block but you have multiple blocks in there. Here for instance we have a function which contains an if condition and that if condition is going to either re reach a reverting branch or it's going to return. So we have AB as arguments, C as a return value and then suddenly we not only have stack layouts that we want to prescribe for individual operations but we also have ST layouts that we want to prescribe for transmissions between blocks or rather block input and output layouts which are also going to have to be generated. So now [clears throat] in order to achieve all this we have a component that we call the stack shuffler which takes a specific input of stack which is a this symbolic representation of stack and then we want to achieve a target stack which has a ordered stack top which is precisely the arguments that the next operation is going to need and it has an unordered lieness. information um informed tail end and this tail end then simply contains the information what we need downstream symbolically after the arguments of the next operation have been consumed to carry on without loss of information. So where does this stack to deep come in? And here I have a small example. So, uh, we have a stack with a V1 on top and then all the way down to V18. And we we might want to change this shape of the stack symbolically to the V18 being on top and the V1 being in the bottom. If we cannot remove any of the stack slots in between because that would cause us a loss of information that we couldn't recover from symbolically then we are in a so-called stack too deep situation where [clears throat] we would have to essentially execute a swap 17 which means I will exchange the V1 with the E8 V18 directly and that as of now doesn't exist in the EVM op codes. So we we can't do it essentially. And the idea of memory spilling is then you know let's park a few of these symbolic values in memory at specific addresses so we can in fact remove them from the stack so we can work around this limitation of only being able to swap around the top 17 elements of the EVM stack. So here is this example um explained if we did have memory spilling and supposedly we have spilled the variables V18 and V17. In order to swap around V1 to V18, we would first pop V18 and V17, which frees us frees us enough space to perform the swap 16, which brings the V18 down to the last slot. And then through m loading v18 and v17 back onto the stack we can achieve the uh shape that we need. So it's it's really as simple as that. Um this is provided that we know that we have to spill v17 and v18 but we also have to figure out that first of all and then also where to put it and that's basically what it is about with the new back end. And then I also alluded to it in the beginning. We have long compile times. So here's a little bit of a historical view on it over a couple of projects. I don't want to go into too much detail, but on the x-axis you can see the last couple releases. So 0820 through 0835 plus the current development version. 0820 was in 2023 if I'm not mistaken. And since then we have improved the performance quite a bit. So up to two or up to four times the performance we had back then. Two to four times. Um still it's too slow though. And uh as this graph already shows different projects have different speedups. So in the end it's all a question of access patterns and how do I optimize what and figuring out where the bottlenecks are. [clears throat] Looking at that a little bit more, if you call soul C, which is the solidity compiler via IR with the optimize flags and then you trace through the program and emit a few timing information, [snorts] then you can see that about 80% of the time is spent in the U optimizer. And this is an uh statistics I made by looking at Sourceify and about compiling about 6,000 contracts on it and just checking you know where where is the time spent. [clears throat] Turns out that the U optimizer essentially has two problems. The first problem is of algorithmic nature which is something that we can deal with different representations which is this back and about and the second problem is that it was never written with proper benchmarking in the back of people's heads. I mean it is an old project. So we are getting an opportunity to do better now with a new back end which actually is easier to do than trying to rewrite what is currently there. So that leaves us with SSA CFG u and SSA and CFG are two nice acronyms that I want to explain a little bit. So SSA stands for static signal assignment form which means I will transform a piece of code. So on the left hand side I have a non SSA version which is y equ= 1 then I set y equal to two and then I set x equal to y and I will translate that into a corresponding sa form where each variable is assigned exactly exactly once. So everything is constant. We have instead of two variables three. So we have to pay a little bit right. We we get more variables but then everything is constant. V1 is constant. V2 is constant. And then V3 is also constantly set to V2. This is a shape of or or a form of representation that is battle proven in many major compilers such as LVM or GCC. [clears throat] And in particular, it enables us to speed up the optimizer by having more efficient access to dev use queries and use dev queries and all that kind of stuff. So here are the algorithmic gains. Then the second acronym is cfg which stands for control flow graph. As initially uh described, not every program is a single block. If you have multiple blocks, you have to reflect that somehow. And one way of doing that is a control flow graph. [clears throat] So on the left hand side, you have a simple if flow which has three different blocks A, B, C. The B block as a condition. C is then the the post after the if. You can do similar things with four loops where you have a cycle between A and B or rather an edge from A to B and then what we would call a back edge from B to A with a loop post C or on the right hand side a little bit more complicated example where you would have something like a for loop which has an additional break condition in it which allows us to break us out from the body. Now we can throw this all together and this is then what the SSA cfg representation of fuel looks like. Um yeah I hope you can read this properly. So we have an entry block and then you will notice that there are not only ordinary variables in here but there are also so-called five functions. As I said before, everything is constant. But if everything is constant, how are we going to do loops? How are we going to set set a value of a variable conditionally? Right? And this needs a sort of trick of functions which are defined on the control flow graph itself. These are these five functions which essentially tell us let's look at an example here. 51 for instance, if the control flow came from block zero, which is the entry block in this case, then the variable 51 is going to have the value zero. And if it came from block three, which is kind of the the back edge from the outer for loop, then the value is going to be whatever the value of 111 is. So like this you still have constant variables everywhere but you have this additional construct of five functions which are these functions on the graph structure of the program itself. Now doing uh the ssacfg representation based on ule allows us to do stack layout generation slightly different than it was done before. So if I hold them side by side, what we do now with this new back end is we go top down. You look at the entry of a function call. You could look at the entry of the program itself and then produce essentially top down really in a topological sort style of way or topological order style of way the ST layouts. We base this on explicit livess analysis. Lifeness analysis tells us at which point of the program is which variable life or needs to live on and which variables can I forget about. This is a crucial information for stack layouts, right? Because I I need to prune stuff away. And then if I have control flow joins for instance when I have a variable whose value is set conditionally then I know by the propagation of information that all information is going to be available at both parents of the block that I looked at and then I can simply pick one for instance. So this is a very constructive way of producing stack layouts which is also going to be good for us because it produces code that is actually maintainable. Compare that to the current vir which doesn't have an SSA form. It uses a bottom up stack layout generation because of that which means you look at function returns you look at reverts and stuff and then say okay at this point in time that's my current stack layout and I'm going to propagate it back up that is done because you don't have an explicit livveness analysis and in particular you then will only see the definition of a variable [clears throat] after you have seen its use. Um this also means if I have control flow joints in the bottom up direction then I cannot assume that I can just pick one of the two but I have to actually merge two stack layouts which is a very expensive and opaque operation. So we're winning there [clears throat] and then essentially to the meat of it. Uh so the current via our pipeline has something that is called the stack limit evader which operates on the u a dual a is being transformed by replacing expressions through m loads and m stores. So the the uh stack to memory spilling happens on an a level. This doesn't have anything to do with the EVM yet really other than you know you being an EVM designed language. So to target EVM but it doesn't really need any of the EVM specifics at that place. Um, if we look at what the SSA form enables us to do, we can simply say, okay, if I do the stack layout generation, I don't even have to touch the a anymore, but I can defer the stack to memory spilling to the point in time at which I actually need to access something to the point in time at at which I actually would have to spill something to memory, which is precisely the stack layout generation, a step after the U optimization [snorts] and in this case it's then going to be on demand. So if the stack shuffler this magic component which transforms one ST layout to the other gets stuck and cannot uh continue anymore to achieve the desired final form. it's going to trigger the spill and the spill then translates to a dupe and an end store at the definition side because we only have one definition of a variable and that definition is constant and that is going to go into memory and and that's it and then later I will push the address and I'm loaded back and this is then done in a loop until we can actually generate the code and there's no more stack leap. So yeah, going back to the initial example, the um S SACFG pipeline, as Jacob already mentioned in his talk, is available under an experimental flag in the latest release. The memory spilling is not yet part of that release. It's going to be part of the next release. It's merged into the current develop version. And if you take this example then you can actually compile. There are of course caveats. Yeah. When whenever we have situations that we think that there are no caveats we will find them anyways. And one of them is that we cannot really handle recursive core chains. If we have a recursive core chain we don't really know where to spill in a finite amount of memory. You would have to bound that somehow. you would need stack frames and such. So we are intentionally saying if we are in a recursive core chain then we cannot do memory spilling. And if you do write inline assembly in your solidity contract then these regions have to be marked as memory safe which means you as the developer promise that you don't mess with the solidity internal memory model. That means you know you you yourself take responsibility to some degree that you don't destroy the internals because if you do then all bets are off essentially. So these are the two caveats. Then [clears throat] the IR itself that's kind of if you're interested for this many return values example looks a bit like that. You have the memory guard which is a property of the entire contract at the very top which is set to some value. That value signifying from where in memory the quote unquote user space begins. So which memory is freely available for use. And then there are instructions. So v 0 is set to the memory guard value itself. V1 as a constant. Then we have a built built-in call to M store and then we call into this uh userdefined function many which is v3 and then return a tupal value which is projected down to the individual components of that tupil. These are these projection instructions and in the end we have a bunch of m sores emitted. Yeah. And if you look at the spill info, you can see that indeed exactly two values have been spilled to memory B5 and V21, which is then as I explained demand driven by the ST layout generation phase. Okay, so what's the current state? I already said it's on develop including memory spilling. We're still ironing out a few kings, but under an experimental flag, you can use it. You can try it out. Uh the bite code performance of this new back end is mostly on par and also surpasses the current VI pipeline at times and the compile times as we haven't begun porting optimizer steps hasn't improved yet. So currently would still run solidity code u IR ule optimizer which is you know the slow bit then transform this optimized duo into the SSFG representation which then triggers codegen. [snorts] Then next up is stabilization of [clears throat] this pipeline. We have the experimental flag for a reason which means you know you can try it but if you want to use it with grown-up money then please don't or you have to make absolutely sure that everything is correct in the end. The stabilization phase is supposed to alleviate that. We want to have thorough code reviews. We are actively doing fuzzing on this pipeline. And um then finally there are few remaining bumps in this stack shuffler which is a greedy algorithm in which it is easy to sometimes get stuck in loops but we are ironing them out. And then finally, and this is a more midterm prospect, we will port over the optimizer steps and actually hopefully get a compiler which is fast. Yeah, that's all I got. Thanks.
