Raluca Diugan - Mechanisms for Unlocking Idle Blobspace
ETHCluj Meetup·Tue, Oct 7, 2025, 12:00 AM
Ethereum's rollup-centric roadmap requires that data is made available on L1. Introduced via EIP-4844 as a cost-effective alternative to calldata, blobs are data objects used by many rollups today to publish their transaction data or state diffs to L1. As the demand for blobs increases, forks such as Pectra are set to boost blobspace supply, thus preventing blob gas fees from becoming impracticable. In tandem with such steps along the Danksharding roadmap, we recognize that blobspace is not always efficiently utilized, with bytes being lost to zero padding or because of uncoordinated commit schedules between rollups. Emerging proposals advocate for more efficient blobspace utilization via mechanisms such as blob sharing or dynamic blob sizes. In this work, we survey such proposals, seeking to understand their feasibility and suitability for different rollup designs, such as zkSharding -- a sharded validity rollup architecture. We aim to consolidate a research direction complementing the Danksharding roadmap, whereby blobs are most efficiently used, maximizing the economic benefits of EIP-4844.
Transcript
Um hi everyone I'm Ruka and I am a protocol researcher at Neil Foundation and today I will be talking about mechanisms for unlocking idle blob space. Um the talk is structured into two broad parts. Uh first we will go over how we use blobsp space today and in the second part we will think about how we can use it more efficiently and um just a short overview of what blobspace even is. Um it was introduced with EIP4844 uh more than one year ago and it was introduced as a mechanism for rollups to publish their data more cheaply to Ethereum. uh because before that the only option for them to do that was using call data which u became expensive with large amounts of data and even more so since spectra um and the blob scheduling algorithm u or the blob scheduling specification defines how many blobs can be used for each fork.
Um all right so there are three important numbers that we actually care about. Um first is the number of bytes per blob which is 131,072. uh and this uh so each blob is structured into 32 byte field elements. The second number we care about is the number of blobs per transaction for VA transaction which is between one and some limit number of blobs. In this case we have a transaction of two blobs and the number of blobs per block which can be between zero and the same limit number of blobs.
Uh in this case we have two transactions uh the first one carrying four blobs and the second one carrying two blobs. Uh and they need to add up u maximum to the limit. Um and this limit as well as a target which is smaller than the limit are defined per fork. Currently for Pectra we have a target of six blobs per block. Um and the limit of nine.
All right. Um let's look at how we actually use blobs today. Um and when I mean blobs I mean uh both in terms of uh the blobs themselves but also the bites. Um so this is a graph from blobcan. And if we look um at this uh graph um if we look at the last part since spectra so since May six or seven uh we see that we use around 27,000 blobs per day and in terms of data the distribution is the same because all blobs have the same size but in terms of data we use around 3 gigabytes of data per day.
What these graphs don't show us however is how much blobs we don't use. Um so let's look a bit at that. Um I looked at the last 15 days of Denon and the first 15 days of Spectrum and um here are some of the numbers. Um whenever I talk about um transactions by the way keep in mind that I use a 180 bytes average per transaction. Um so um for the two halves of the data set um I talk about the target um because it's where the equilibrium between supply and demand uh so the blob space supplied by Ethereum and the demand uh asked for by the rollups um is reached.
Uh so for Denon we had a target of three blobs per block and for we have six. Uh in terms of provision uh blobs, how many we could have actually used w within this 15-day period perfork, we have around 300 for Denkun and uh double of that for Pectra. Um and the number of an average uh number of blobs per block we actually used comes very close for Denon with the uh target whereas for Pectra we have very we still have room to scale up to the target. So for Denon there was very little space to scale by adding more TPS. Uh and in terms of the blobs we have not used um this number is insignificant for Denonkan and it's around 40% of the provisioned 650,000 um uh for Pectra and in terms of TPS that we could have used uh it's less than one for Denkun which again means that we could have we couldn't scale further and for PCRA uh it's around 143 meaning that we could still scale currently um rollups um to add to the 200 160 or so TPS that they use collectively.
Um and if we look at how this uh is distributed, we can start with a blank canvas. So here we look at 100 blocks uh first for Denon. Um and the gray area is the target. Uh so if we used um all the 100 um uh blocks up to the target of blobs, it would look like the gray area. Uh and this is how it actually looks like in practice.
For the last 100 blocks of Denkon, we see two things. First, um there's quite some blocks that go over target uh in the in orange. Um all those go go uh beyond the uh target of three blobs per block, but also there um blocks that include no blobs, the gray uh spaces. Um so this already hints that maybe we could do something with orange parts and move them a bit into the gray area for Pectra. Um things, sorry.
uh things look even more relaxed. We have only a few uh blocks that exceed target and we have quite a lot of uh gray blocks which uh store no blobs. Um and so this was in terms of blobs uh but also uh we can talk in terms of bytes how we use blob space. And remember that one blob has 131,72 bytes. Uh however, this is an example of one blob that only uses around 150 bytes of those.
Um and this is actually an almost 100% waste rate. All of these byes store no meaningful data. It's just zero padding. Um, one study actually uh looked at the first six months of Denan um and found that the data size that we put into blobs uh goes from like 0.14% up to 99.
99% of the actual blob. Uh but if we flip this number, we actually get a 0.01% to 99.86% waist rate, which means some rollups waste up to that uh amount of bytes. Um, and this points to the conclusion that maybe blobs are not a one-sizefits-all solution.
Some rollups may benefit more than others. In particular, small rollups um have uh the following two options. So, uh small rollups or uh rollups that output uh lower amounts of data uh have the options to submit uh partially empty blobs. So they just pay for zero bytes with no meaningful data or um they can delay their DA submission until they collect enough um data to um post a full blob. But this makes for a bad user experience because users would have to wait until um the rollup state can be finalized.
Okay. Um and yeah, thinking back at the uh data set that we um studied with the um last 15 days of Denan and the first 15 days of Spectra, for each fork, we wasted around 5 billion bytes. So these we actually paid for uh we completely paid for them uh by publishing them on Ethereum. Uh but they were used um to store nothing. uh which translates to 30 million around 30 million transactions um and two million transactions per day.
And if we translate this to TPS, this is between 22 and 23 TPS uh that we missed. And this is actually close to uh the number of TPS of the second largest rollup today. And this means that we could have actually stored an the data of another uh second large rollup today for free by distributing all of its transactions across all the different blobs where we have zero bytes. Okay. And just to look at where this uh hidden wasted bytes are, well not hidden but where they are.
remember the first um 100 blocks of PCRA and here you can see exactly uh where the wasted by bites happen. For example, here in the middle um there is one block that has uh waste over target but if we remove that waste it actually the the block itself goes under target. And going back to our title of idle blob space um I would like to use the words of someone famous to define it. Uh it is not only idle the bite which does nothing but it is idle the blob which might be better employed. And in our example these byes are idle because they store nothing.
And these bytes are these blobs provisioned blobs not used blobs are idle because they could have actually been used better by storing uh those other blobs from a previous block. All right. So now we're moving on to the second half of the presentation on how to actually u try to make it more efficient use of um this blob space. Uh and there are a couple of approaches. We could adjust the data.
Uh we could adjust the blob, we could adjust the the block or we could adjust the transaction. And we'll we'll go over a couple of examples. Um yeah, next. Uh so first we could adjust the data. Uh for example, we could decouple the batch size from the blob size.
And this doesn't apply to all rollups today. Some already do this. Uh but some don't. Um so the way some uh rollups batch their um or define their batch is by um limiting the amount of data that they put into the batch by the blob size. Uh which results in having this configuration of partially empty blobs.
Um, and for example, the transaction that I linked here uh contains three blobs. Um, but 26,000 of those bytes are wasted and they could have actually stored 143 more transactions uh if they stream the data as it shows here. So all the empty uh collected uh byes at the end could have been used to store more transactions. Um the next approach uh would be to adjust the blob. Uh one um approach in this sense is um to use dynamic blobs.
And the idea the idea here is to um limit the blob size uh by the data size or to define the blob size by the data sizeish because we would still use some padding but much less than uh in some cases today. And we can take an example of some 88 bytes of rollup data distributed over 32 by fields field elements and um have something like this. And then we realize we need need to complete the last um field element and we just add it with eight bytes which is much less than um a third I don't know 130,000. Um and this approach uh in this approach we would um define we could define the blob space per block as the total number of bytes instead of how it is currently defined uh as the number of blobs. uh and also the blob would have field element granularity meaning um each the size of each blob would have a 30 would be a multiplier of 32.
Uh one benefit of this approach um if you think a bit more game theoretically and who would be interested in uh following this approach is that um Ethereum gets the full execution cost uh paid for because all rollups still pay for the service access but they only pay um for data proportionally to uh what they actually used instead of uh paying for more data that they need. And uh if we're thinking about how this is going to be compatible with u future upgrades. So currently Ethereum is looking into introducing sambling um and if we modify the size of the blob then we we should also think about how um to make sure that all blobs have uh efficient uh sufficient replicability and one proposal looks into this um and they propose to use multiple sample sizes um to achieve that. All right. Um so another approach is also to adjust the blob uh but not in terms of size but in terms of content.
Um and this is uh blob sharing. It's a very straightforward approach. You have multiple rollups. They have multiple um each of them has multiple data shares of different sizes coming in at different intervals. Uh and at some point there is a blob sharing service which collects all of these data shares.
Um it applies some aggregation algorithm for example first come first served or the largest um uh data share first or whatever. Um and then it applies some transaction submission strategy to determine when to post the transaction then um so the transaction is ready and it's posted to it's sent to Ethereum to a shared contract from where all the other labs contracts can read from um the data that is relevant to them or the information. Yeah. Um one question in the case of blob sharing is who actually operates this uh uh blob sharing service. Um so u two possible two possible approaches for the operators is the rollups themselves.
Specifically um in case of um clusters of rollups which have already some common core um around their stack they could easily adjust this and they could do it uh fee free. Um and the other option would be to operate uh to have it operated by some third party service uh which might charge rollups some fee for this. Um yeah and these are some of the um options that already exist at different stages of development. Uh some are just prototypes, some are actually in production um if you want to study them more. And one of the studies that I mentioned earlier which looked at the first six months of data um also looked at um the benefits of um sharing blobs uh particularly for small rollups.
So the roll-ups that have small amounts of data and first they did notice a slight increase of data from uh the added over overhead because rollups do need to be able to identify their data shares within a shared blob. However, and obviously they also noticed a decrease in the number of blobs because now all of the zero bytes were actually employed. They were doing something useful by storing uh other rollups data. Um and yeah, the total number of blobs was lower. Uh which is beneficial for for the entire ecosystem because more resource is available.
All right. And so at um one example of blob sharing is what we do at Neil. Um so at Neil we build zk sharding which is a validity rollup architecture where data and execution are sharded. Um and we have two broad options for doing state updates. So this is um three different shards which do cross shard communication between them.
Um and we could either individually update this update their state to Ethereum or we could that do that in aggregate. Uh and we decided to do so in aggregate. Uh and for that we have a synchronization layer before updating the state to Ethereum. Uh this synchronization layer has a module called u the aggregator which collects data from all the different shards. Um and then it applies some aggregation algorithm which is much easier for us.
It's just first come first serve. um or alpha numeric ordering based on charts because it's the same network. Um then the batch committer takes this um data, prepares a transaction, posts it to Ethereum. At the same time, uh the prover gets this data and uh provides generates a validity proof um sends it to the state proposer and um then the Ethereum. Um all right.
So now coming to the last approach we will explore today. Um we could also adjust the transaction and the block. um by building blocks based on the target uh number of blobs. So um the this approach particularly is helpful in high congestion intervals. We're not that there yet with PCRA uh but we were there with um Denkun and um in this case uh in this situation it is helpful because it helps us uh long-term keep uh fees lower.
Uh and the idea is to try and get the target number of blobs for for each block where we have blobs. Um and to do so we could um in order to like get to the target and build it sort of like a Tetris um try to always uh get the first in fitting first out. So we could skip the blobs the blob transactions that don't have the the fitting number of blobs and delay them. Um but in order to uh avoid this delays, we could actually lower the blob count per transactions uh per transaction. So instead of um having like five or six blobs for each transaction, we could have a lower number so that the transaction gets um fit in uh faster.
Um so in this example, uh the first transaction, the first block here actually had two transactions transactions of five and four blobs each. Um but if uh we want to do this approach and only keep it um to the target blobs per uh per block, it would be we would need actually six blobs in the first block and we would have to wait until we get that one blob transaction. Uh but if the four blob transaction would have been split into multiple uh transactions um then we could easily just fit them as they come. Um and there is research that supports that uh rollups will respond to the market conditions by adjusting their blob count per transaction. So maybe this is not an unrealistic um demand.
Uh also there is some other research pointing out uh that frequent smaller transactions do have lower delays. Um and the the benefit of this approach is that if we um try to always get the target loss um to like never exceed the target, we uh don't get into fee increasing territory, meaning we keep fees as low as possible. Um and the downside or major downside for this is that high throughput rollups will probably not agree because uh for each DA transaction they have to pay execution costs. Um so the more DA transactions they do um the more execution cost they pay. So for them it's more um efficient cost efficient to u pack as many blobs as possible.
All right um so this was the last mechanism and now just some uh thoughts to wrap up on how to make dung sharding which is scaling um DA for rollups how to make it more uh sustainable. So currently the way we uh scale DA is we increase blob count with um with every fork and we will also introduce the scaling by we I mean Ethereum. Um but we could also look at this approaches uh to minimize waste and uh some possible next steps into trying to uh introduce this approach approaches into the bank sharding road map uh is to analyze exactly how different mechanisms uh what what cost implications these different mechanisms have um align incentives across different actors from rollups to third party services to Ethereum R&D um and to also think about how they integrate with future um upgrades such as D asling um right so that was me um and I appreciate your attention and I'm happy to take questions Yeah, we have two questions for you. I can read them out or you can just answer them.
Yeah. Um, okay. So, is there coordination underway between rollup teams to pilot blob sharing? Uh, not that I know of. Um, my impression is that blob sharing kind of will come first and then the blob sharing services will try to get u users.
Um but yeah as I said actually cluster rollups uh might benefit from it if um a lot of like uh small rollups exist uh built around um like this um some common core uh it would be meaningful for them probably to to try to do that. Um I hope I answered the question. Um will block space increase to hold lampert signatures? I am not sure if this question is necessarily related and I'm not sure I know enough about Lamford signatures to answer that question. So
sorry, can I just pass you the mic?
It's giant chunks of hashes for postquantum signatures. So it's a lot of data.
Okay. Um, and you mean Ethereum block space or blob space? Um
I Okay, so my impression is that okay I'm not sure what the amount of data you mean but um if we're talking about generally what Ethereum is aiming at right now is I hope I'm not wrong but around 256 blobs per block uh times uh 131,000 bytes. Do you think that's enough? Um Right now there's smart contracts that use like POA and so they basically have to put the lamport signature offchain and then they have to like prove into the network but that is having a second party that's being basically trusted. So if you could put half of the transaction into the blob and piece of it referencing from from the block where the the main transaction data would normally be held then the signature verification would be stuck in the blob which would change the execution of the way that the block and blob works.
Okay. I am I don't think I have the answer to that question. I'm sorry. Do we have any more questions? Oh, is this still for us?
Oh, no. Not for us. Oh, good. It's good. Never mind.
Uh oh, I think we have one more now.
Yeah.
Okay. Uh, do you see a future where blob space becomes Oh, okay. Okay. Uh, do you see a feature where blob space becomes a resource that can be programmatically traded or auctioned between rollups? Um, so it is a possibility.
How realistic? I'm not sure. Um actually part of the like blob sharing approach when I say rollups can be the blob blob aggregation service. Uh it could also be that one rollup just auctions off u some um like remaining bytes of their um of their blob. Um
so it could be uh becoming more advanced than that. I'm not sure if the effort kind of justifies uh is justified. I don't know yet the exact cost implications of that. Uh but yeah, like theoretically it is possible. Um yeah.
Um are there any blobs used exclusively for NFD transactions within the VVM or not yet? Um I don't think there it's an easy way to find that unless there is just one roll up that is specialized on that and they use blobs that would be an easy way to check. Uh but you can post other random things for on blobs for sure. Um I've seen some examples where not roll up data is what's posted. So yeah, you can use it to store music for 18 days if you're interested in that.
Automatic transcript — names and jargon may be misspelled.