New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Scaling Ethereum L1 with PeerDAS - dapplion

ETH WarsawTue, Oct 7, 2025, 12:00 AM

Ethereum scalability roadmap is focused on increasing its data throughput without sacrificing on decentralization. After many years of research, Danksharding has emerged as the clear way forward. Now, Ethereum core devs are implementing its first iteration: PeerDAS. PeerDAS re-uses battle-tested p2p components already in production in Ethereum to bring additional DA scale beyond that of 4844 while keeping the minimum amount of work of honest nodes. 🧜🏻‍♀️ ETHWarsaw is a series of educational and entertaining events for an active community of blockchain builders, developers and enthusiasts with focus on Ethereum-related tech. Once a year, we organize a large conference and hackathon for the community in the center of the Polish capital with speakers from the best web3 projects and participants from all over the world. Follow ETHWarsaw on social media for the latest updates! X (Twitter): https://twitter.com/ETHWarsaw LinkedIn: https://www.linkedin.com/company/ethwarsaw Telegram chat: https://t.me/joinethwarsaw See you all at our events in Warsaw 🙌🏻

Transcript

thank you for the wonderful introduction and uh happy to be here it's a pleasure first time in Poland so far like it sh so yeah today we're going to talk about scalability and ethereum and just to get everyone on the same page I'm going to try to accomodate the talk both for like people that don't have that much experience with the topic and experience so when we talk about scaling we talk about the capacity the chain to have as much transactions as possible at the price that we them reasonable so say if ethereum today has 10 transactions per second we want to scale we want to elevate that to 100 to a th000 more capacity means cheaper fees and at the end that means that we make the chain more democratic more people can access it now ethereum is like hyper financialized defi only type of usage if we can scale ethereum by now I don't know 10,000 we could have much more diverse typee of activity more things like social gaming Etc and accommodate other types of um communities that maybe don't have the resources to pay the transaction fees that we have today so as you were hinting the whole thing is this scalability trma so we could scale blockchains and some blockchains have choose to do that today but they usually sacrifice something the easiest one is decentralization you just have chains with very few nodes that are very expensive to run and they are run by a few cartel of people and use sacrifice on decentralization but in etherium since Genesis the key has been to preserve the transition that's what I think would make ethereum thrive in the long term so to motivate what I'm talking about and give you a precise example let's understand why we are going with the rollup Centric road mapap so consider our two beloved chains Bitcoin and ethereum where again for this installation we want for everyone to be able to run a note so we have this red line we want to make running nodes cheap and affordable so you can have them at your house with modest hardware and a reasonable internet connection what that means is that we limit the throughput to the chain to the lowest spec participant now if we want to scale we can do some example of a chain I'm not going to mention where we just make the cost of running the not more expensive more throughput but we lose this capacity of everyone to run nodes then the chain may caliz we have maybe I don't know 10 BSC Farms running the nodes they can start to do things they can start to censor people they can do irregular State transition they can do things that are damaging and potentially destructive to the chain and the the large base of users don't have any records because they don't run nodes what they have to do is they have to go to Twitter and be like uh you have done something bad no then recover which has happened in the past we did it in ethereum it was the dard fork but it's very costly it's not optimal takes a lot of Social Capital reputation it's not something we want to build ethereum on so what if we could automate have this very expensive social recovery step done automatically with math and that's exactly what rollups do we upload execution onto this layer twos and then they with these clever tricks can prove to L1 that they have followed all the rules automatically and if they don't do we have this recovery done automatically so optimistic rollups they do it with economic games and dissection games and valid rollups with crazy cryptography but the end is the same execution happens there we prove to L1 that's correct done now that we have addressed execution the next point is data so what happens if data becomes unavailable and with data I mean the transactions that go into the rollup or the state divs that go into the validity rollup if data is not available we are screwed so for optimistic rollups we have to do this bsection game and if you don't know what's the input state to the optimistic rollup no one can dispute the transactions so basically if you submit an update say to optimism and you don't provide the input transactions you take complete control over optimism and you can steal all the fans not ideal for valid rollups it's not that bad but you can transition the rollup into a state that only the attacker knows how to progress so basically I would take I don't know ZK sync progress to a state that I don't know and then no one can move their funds unless I do so I will go to you and say hey give me 50% or I don't give your money so not that bad but still um completely destructive so now the job of L1 is guaranteeing data availability and who has to do this guarantee who has to do these checks who has to check that data is actually available everyone that's the main problem that's why this is such a hard problem and this is kind of the last stage on blockchain scalability that we are triggering so to make you understand why this is the case why everyone has to check it I want to explain this case if you have never seen this slide before please please pay attention this is absolutely critical to understand why blockchains go the way they go so execution that we're talking before is just math you have 1 + 1 equals 2 it's something you can prove something you can show to L1 that has been done correctly or incorrectly however the existence of data is subjective it it cannot be proven so let's go to the slide consider case one we have two roles we have the publisher this guy and we have the slasher the one that's going to like police make sure that everyone publishes data which is this one in case one the publisher is malicious and does not publish all the data but then the slasher correctly identifies this fact and says hey I'm not not all the data is here alarm blah blah blah now for case two the publisher is honest and publishes all the data but now the slasher is malicious and falsely raises the alarm now you as an observer come at T3 at the end and you see this two cases you cannot tell the difference in case one there is an alarm and all the data in case two there is an alarm and all the data who's that fault publisher slasher we don't know that's why yeah fact the Publishers we because you canot know uh you cannot build like a an incentive structure on top of this not you don't know so yeah to recap execution we scale with rollups execution happen somewhere else we prove the result on chain and data availability everyone has to check it but we will scale it with data availability sampling that's the next topic we're going to check so everyone has to do theability checks yes but there is a smart way that you can do these checks in a way that scales so the naive approach what's happening today in ethereum and all blockchains is everyone downloads everything so the data is inside the blocks everyone gets the blocks so everyone checks everything this is very safe but it doesn't scale the next idea a bit naive would be just check some of the data just go on the blog and say I'm going to check like this bit and that bit and that bit and if it's mostly available then I guess it's mostly available and blah blah blah we continue that's that's scales of course but it's not safe why if only one single transaction or even one bite of one transaction disappears those terrible attacks the ones that you can steal everything from optimism and sing all that still applies so you need a way to ensure that all the data like 100% of the data has been list and that's why we come to Ure code extension and this is kind of the key magic that makes everything possible instead of sampling the data itself we will extend it with a polinomial so if now we have only 50% of the data we can just compute this polinomial and extract the data back so now the question is not all the data is available but 50% of it and this schol properties now I'm going to just pick some sample and if these samples are available then we have a probability of two over the number of samples I have chosen that 50% of the data is available math but in practice what it means is this now like with probity exampling if you check 10% then you have 90% chance that you will be full that it's actually unavailable but with the extension it's practically zero it tends to zero very quickly if you have enough samples so which is great you know you see check a little bit never fol excellent so this new data that I'm talking like these extensions and polinomial all of this it's already on ethereum it's already in chain with this new space that we call blobs so now we have block space for transactions and blob space for the transactions that the company disel transactions and that's where rer put the data the cool thing about detaching the data into something else is that now we can do different things with these blobs we can expire them we can make them cheaper because blocks are kept Forever on the clients so you need a lot of gigabytes but the blobs after two weeks they're gone so we can have much more throughput without sacrificing on home stakers capacity to store data and the other cool thing is that blobs have their own fee Market so we have one gas price for block data um like the gas price the one that you know and love but blobs they have their own separate gas market and so far that has proven to be very useful uh I'll show you now some graphs and it has resulted that blobs has been free basically since launch which is quite cool so so far all of this is like the long-term road map of what's going to happen with data ability sampling blah blah and everything but a couple like two or three years ago rabs were already like scaling and going to production they were like come on give us data you know we need we need something now and all of this datability sampling is very complicated we haven't figured that out yet so we did this b0 which is called Proto sharding where we do the crypto part and the blob separate fee market and blob space first so we can deliver already this is already in production since March and rabs are happy so what has been the adoption so far I would say pretty good so we did the fork on March here we had blog scriptions not sure if you remember but that kind of blew up and then the market has been flatlining on the top and we have like adoption from all these very cool rollups and because they don't actually reach the maximum which should be around here the fee is actually zero for most of the time or I don't know one to e to the minus 18 something like close to zero which is very cool because if now you want to transact on base or optimism the fees are zero I guess like not exactly zero for reasons but practically zero which is cool it's maybe not going to be like this forever but it's nice it's the first time since forever that ethereum is a little bit under capacity and this this original goal of democratization and allowing different types of activity it's kind of coming to our reality which is really exciting and bullish in my opinion however some people think otherwise so because fees are low now we are not burning eth and eth has been inflating again so I wna I want to do a question for you guys again what happened here is we scaled ethereum fees went down so inflation is up the fact that ethereum is inflation for low fees is that a bad thing raise hands is that a bad thing or is that a good thing okay there is no right and wrong answer like this is economic policy um I think yeah maybe in my in my honest opinion if I had to say I think this is a good thing like ethereum inflation is fine like we should prioritize the fact that etherum is useful to as many people as possible than to reward holders but again that's my take cool so now that we have protot and sharding on mainnet the next step is to actually do sharding that's what we're working on and you can understand sharding with this ratio of how much capacity exist on the Chain versus how much capacity you as a node Runner have to download and process at home now we are here no sharding so you have to do everything but as we move down the road map you will have to do less and less and less but getting to this High degrees it's really really technically difficult so the first step which is what we're going to talk today is pias that's some reasonably modest sharding Factor but still quite quite good because remember we are not at capacity so if we double we'll be like we'll go above the roof just with a 2x so rollups need I don't know some way to catch up and just onboard more users whatever and yeah to recap why is sharting hard so when you shart now you need to find and retrieve this data from parties that you don't know where they are so there needs need to be a lot of network constructions to find those peers retrieve this data and do some operations that are complicated and expensive like reconstruction we're going to talk that later but yeah doing that quickly enough for blocks to become canonical in the attestation deadline which is 4 seconds that's pretty pretty complicated and we haven't figured that out yet so we can do that now safely with a 2X but for like 60 4X we need crazy Network constructions that we haven't figure out yet how to do so to give you some visualization how this is going to look today we have these blobs and we will represent them with rows so this just data data data we don't care just random data from rabs now with Shing 1D P we will take these rows and we calculate the extensions the polinomial extensions we're talking about this is this orange line and now we chop them ver IC Ally into columns so a column is the unit that we're going to move in the network both for custody and sampling and it includes uh a chunk of every single blob or it could be the extension doesn't matter remember that we have to sample the full space extension and non-extension so that if we get 50% then we can run reconstruction and get the blue one which is what matters to us so how it's going to work we have the proposer oops we have the proposer who's going to compute the extension this thing we saw before this is some relatively expensive cryptography then it's going to chop the map into columns and it's going to know which custodians expect which columns so it's going to forward them publish through gossips app that's the first step and this is called subnet do we'll explain it now and then the second step is actual pieras so this group of Samplers sometime after will discover these custodians via dp5 and query them over RRP these terms don't matter much but what's important about them is that these are not new network constructions these are protocols that we have been using on production for four years now we understand them very well we have implementations they are safe it's just that the way they are used here is Noble and because of this novelty the problem now that we find is timing they work but they are a bit too slow so you as an attester of the beon chain of ethereum your key role is to boote on what block is the head and what checkpoint you want to justify and you have to do that 4 seconds into the slot that's quite short especially because today with me and I'm not sure if you heard about timing game names the proposer has an incentive to publish as later as possible so I think now the median so if this is zero now the proposer is actually here proposing here when it should do here but they are greedy and the thing is the later you propose the more Meb you capture so proposers have an incentive to delay publishing and that's kind of messing up this window so if we expect participants to sample within this deadline what's going to happen is that some people will be left behind and guess who's going to be left behind it's going to be the solo stakers the people with the least resources will be penalized will vote on the correct incorrect head and will have an economic penalty so this would become a centralization factor and we don't want that so what we come to the realization well especially Franchesco he's a very smart researcher at thef uh is that we actually may have a better option so to recap in sampling and data ability sampling we need to provide these two Key Properties so this one is the most important which is global safety we want to ensure for rollups that blocks are available with very high probability or in other words that a block that's not available will never become canonical and finalize so that at most as small minority of sampling nodes can be tricked that's very easy to achieve because it's just a probabilistic statement if you have enough of this getting enough columns at random sections just for like sheer overlap you can provide this guarantee then the one that's more hard is individual safety so it's you on your own point of view can you be tricked by a malicious proposer into believing that something is available when it's not and that's very hard to do because imagine that you are the proposer and you know that the sampler note number 15 wants columns number 1 2 3 4 so you just get these columns and give hey here you are these columns and then you hold the rest then the node would believe that it's available when it's not so in order to achieve this one we need a way for Samplers to be more Anonymous or to be precise to have the queries that they have be unlinkable and ideally distributed amongst different participants but with with this initial step which is subnet do that's very hard to do because it's very public so what we are going to do is we're going to divide these two sections and we're going to say we sacrifice uh like the second property so we move the attestation deadline forward only to the first step of distribution to the custodians so that way uh we don't have this penalty that if you have slow sampling you can still attest to the correct thing the problem now is is that this attesting bat doesn't have as much weight so it's a bit more unsafe so what we will do is make these guys hold much more data so we we still have a safe construction but the sharding coefficient may go from four to two but it's that that's okay we can iterate later uh so yeah to compare subnet sampling this is the initial distribution we do over gossips up it's relatively efficient but it's quite expensive in terms of bandwidth because due to how the protocol work data gets distributed eight times it its value and it's also extremely easy to to trick a note into believing that something is available because you have to announce ahead of time hey I need column 1 2 3 4 inste P sampling is much not a lot but reasonably safer in this regard because you don't have to announce aead of time what columns you want and it's perfectly band with efficient because you just do one Quest and you get one response the problem is that this is much slower so we're going to probably do a construction like this in the future where we sample but we take more time like six minutes and then this signal will not be used for a testing but it will be used for transaction confirmation block proposal and justification so for now I think the first iteration of pias we would do cost distribution and on so this is much easier to pull off uh in initial death Nets we have seen instability with sampling so if we just do subnet Dash we can deliver this Center and do the 2x now and then when we figure this out we do the 4X and we just keep going so so far the progress uh we did so this was a big uh Milestone we went all the core deps working on this to Kenya and had a bit of like a hackaton weekl long hackaton where we significantly progress on on pias and got the first definite then we did the second definite on July which died very quickly then we did the second definite on August which is still alive and hopefully by Q4 we'll have production implementations we'll see maybe with P sampling or not still undecided and then hopefully uh yeah this is very optimistic I would say 23 q1 let's put it that way the pectra hard fork with uh pias so just some random metrics to show that I'm not bluffing we are actually doing this we have nodes today running that uh I think it's still too early to measure in terms of pure bandwidth what's the actual factor that we getting but so far it's looking good I'm optimistic clients are getting quite productionize now and uh yeah we'll see in the next four months will be ready so yeah that's beas now road map next steps is we have to we'll keep going we're not going to stop if we get to here we'll have this uh end game that balic wants to call like the 32 megabyte block where we can have like tens of thousands of transactions per second with ton of rollups but to get there we need to do this and this is much more complicated uh because what's what's going to here is that to get to this very high sharting level we need the load on the single note to be really really really small compared to the amount of data so I mean I don't we don't have time to go over this but the key idea when you'll see this picture that you may see in the future is that instead of sampling the whole column now you're sampling a single square and this is going to be really big like we're going to have this is going to be huge so this cell is going to be tiny just kilobytes of data and if you just check like 32 of them you can prove to yourself that the data is available and that's the that's the magic of this construction that you have a huge amount of data and you just available both and on So yeah thank you so much hope it was not too reing think we have times for some questions yeah you said there was a probability [Music] 0 yeah where you had listed there is a z yeah I think after that yeah yeah here so I mean considering the amount of economic value that is stored in blockchains if it attribute an economic value to this that is still significant right the point that I'm trying to make here is I mean terrific great approach loved it totally but someone still needs to solve for long-term data availability and you know where everyone downloads everything but in a more infrastructurally efficient manner I'm not sure uh I'm not in this field but I heard like some covalent is building an ethereum wayb back machine or something so is that also a a need for the industry right so that's a that's a great point so the whole point to provide this level of scalability is to make the blob data expirable and that's That's essential it doesn't mean that other types of services or industry can take on the role of keeping that data and that's what's already happening if you want to access all blob data today because it has already been inspire you can go to Ether scan or blob scan or services like this that they would give you the data if you want but it's not the role of the protocol anymore and we have to do that otherwise it's not going to uh scale um if you I mean hopefully not but if you so what what he's mentioning is that ether scan may also remove the data and I don't think they will because they have an incentive to keep the data and it's not that expensive to keep the data the the key point is that we are removing this this responsibility from testers from the validators so you have to think on the hom Staker now the home Staker don't have to keep this data anymore it can forget about it so it now it's it's the role of a different type of entity and if you are a rollup or an exchange in other means someone that's really really interested in having this number with 99 zos then just keep all the data yourself there is no one yeah there is no one that's preventing you from doing that or this number could be much bigger if you just sample a lot more or if you're really concerned just get all the data is that part of ether technology or it's your own chain sorry is that part of ether technology or it's your own chain I mean it that that solution will be include in near future in ethereum or it's something a different project yeah everything we talk today is on ethereum mayet no but you say mayet so I'm asking about that solution that you present so uh the the future sharding it will be part of ethereum right yes and you you think it's really possible that in next year it will be uh in mained yes P do next year mayet I would say 90% chance or 100% any more questions no so difficult sorry thank you thank you thank you

Automatic transcript — names and jargon may be misspelled.