Sszb: A High Performance SSZ Implementation in Rust
Devcon·Tue, Oct 7, 2025, 12:00 AM
This talk goes over my EPF project for the SSZ ecosystem: - a benchmarking suite for the various rust SSZ implementations in the ecosystem to properly evaluate performance and point developers to which library they should use. - a high performance ssz implementation that's faster than existing libraries in the ecosystem
Transcript
[Music] um hello my name is gilia I'm here to present my work on SS ZB a high performance SZ implementation in Rust as well as my work on ssz Arena a benchmarking suite for the crates in the rust ecosystem um my motivation for working on this project project was uh having a chance to work on optimizing an existing project uh which I haven't done before and so I learned a lot in doing this as well as um seeing what kind of gains were possible in something like ssz it's not a real bottleneck in uh clients today it's still a a fairly simple task but I was curious about what kind of losses were present over a long period of time uh if without an optimized implementation and so I thought this work was uh very uh appropriate for the fellowship because it's I guess lower priority for client teams but still work that needs to be done so um I'm briefly going to going to talk about what serialization is and ssz and go over the ssz ecosystem as well as my work on the benchmarking suite and the implementation finally talking about performance and next steps so zooming past these first few slides serialization is a process of transforming a data structure into a common format that uh the different clients can agree on which is uh important for consensus uh normally data structures can't be transmitted as is because of different in memory representations and such ssz is the serialization scheme on ethereum 2 meant to replace rlp on the execution layer uh it has a few improvements like uh schemas which are a must have in a performance system and merization is the process in which we generate a short uh digest of the state um while allowing for updates without rehashing the entire state uh this is useful to generate small proofs of uh the contents inside a beacon blog for like clients that don't want to hold too much State data so uh there are a few ssz implementations uh in the ecosystem mainly and go uh fast ssz is uh most popular one and used in gu today although Peter also has an implementation out uh in Russ specifically there's Sigma Prime's ethereum ssz grandin's crate which is in public and is only used internally and Alex stokes's crate as a zrs uh for the first half of the fellowship I worked on a benchmarking seat to evaluate how these libraries perform against each other uh most of these crates have tests on consensus spec tests uh but there's no real way to see how to perform on real data like Beacon blocks and Beacon States uh so I worked on a uh Library oh not a library a project called as AZ Arena which uh benchmarks the different crates in the ecosystem and so it evaluates it on more control test cases like lists of uh integers and validators validator structs and also evaluates it on blockchain data like Beacon blocks and Beacon state that I obtained from a uh Beacon chain checkpoint so I leverage Criterion for the this Benchmark Suite which gives us handy reports but I also use uh another uh benchmarking Library called the Devon which handily gives us allocation stats so you can see how much memory is being allocated during these uh encoding and decoding runs and so the this gives us a robust way to measure performance now on to the second half of my work into Fellowship working on my own implementation so how does one optimize ssse optimizing this uh serialization scheme and benchmarking is kind of tricky because most the real bottleneck or most of the optimization in encoding and decoding is simply using a more optimal data structure there's a lot of techniques you can do to optimize this you can lay out your data and memory so uh it's aligned to the word boundaries and there other uh techniques like zero copy deserialization where you can simply cast your btes into your type that's only really possible when you have control of the under underlying data structures um which I did not have for this project and there's a lot of reasons why you might not want to just rewrite your types with serialization in mind um Sigma Prime particular has done a lot of good work with their mhouse crate which um allows for faster uh sparse updates of Beacon State and so it doesn't really make sense to just change your type that crate just to speed up serialization it's better to just think about how to work with uh the types we have and so being constrained by the inability to change the underlying data structure I opted to minimize intermediate allocations and so another a second bottleneck in serialization is how much memory are you allocating in between steps to serialize and deserialize and so here's how I went about my implementation as a ZB has two main differences it uh uses the buff and buff mute traits this is an abstraction over uh buffer types so it encapsulates both vectors and slices and has the added benefit of abstracting offset en counting which greatly simplifies the implementation so for context you can outright Define how to encode and decode certain types but the SSB package also provides a macro for automatically generating implementations for container types which are like structs um and so generating this these implementations is a hassle we um provide a way to do this automatically and the implementation for it is very simple thanks to the buff mut trates and second we avoid a lot of intermediate State during the encoding process and minimize any needed allocations during the decoding steps uh this reduces the number and size of memory allocations needed to perform serialization which is another dominant cost as I mentioned before um Peters go go implementation uh SS implementation does this and grandine was also another big inspiration although to note grandine only works with slices uh the benefit of using buff and buff mute is being able to use vectors and any other buffer implementation you want to provide as long as you implement the trait which not quite sure how it's going to be used just yet but could be handy and it performs pretty well so I tested this on beacon block decoding and encoding it's pretty pretty fast on uh the decoding part if you'll consult the graph I'm not quite sure how visible it is but um we clocking in at around 129 mil uh micros seconds on the decoding part versus 3 milliseconds um in ethereum ssz and while the differences aren't as drastic for uh all types we're getting uh there's similar levels of performance um around 85% uh encoding speed up and 95 decoding speed up on the beacon blocks um so that that was for the beacon blocks there's still some changes to be made some fixes to be made for uh Beacon State encoding and decoding I know the I there's a bug in the implementation I know where it is I'm going to go fix it but I only found it like two hours ago so uh didn't really have time to fix that today uh as for next steps I want to ship a a support for merization and Merkel proofs with generalized indices this is needed to have a full-fledged as a the implementation my focus for this Fellowship was on performance of encoding decoding and so I left this for after the cohort additionally I'd like to support uh a new trait I call SS check which provides uh early input validation for to check that a an input conforms to a certain type this would be useful if you want to reject malformed inputs earlier instead of having a full decoding step in the hot path of your application uh and then after that stable release I want to gear up for right stable release adding usage docks and cleaning up anything that needs to be uh Polished in the library and then once that's done I want to work on something I find interesting but I'm not sure if other projects would want to use this but I think it'd be cool to have support for partial encoding and uh decoding for example for large objects like Beacon State fully decoding can be very expensive and so partial decoding would drastically speed things up especially if you only need a subfield of your beacon State and uh re-encoding and rehashing would work similarly again with uh the sigma Prime's Milhouse implementation they're already implementing dir types with uh sparse updates in mind and I think uh ssse could use uh some similar ideas with regards to sparse updates I'm a little over time uh I want to thank my uh Mentor Michael spra from Sigma Prime who did uh most of the work I believe on uh ethereum ssse he's not a def con uh I think um but if you're watching this thank you and I also want to thank uh Josh and Mario for um providing the opportunity I learned a lot through the Cort and uh I'm glad I got the chance to do this work um any questions that's all for me yeah all right any questions about the ssz library is it ready for C to swi over uh not right away probably after stable release I forgot to mention also that there's no unsafe code in this so it's uh my Michael told me not to use that uh yeah um coming soon yep coming soon all right one more time for gilia [Applause]
Automatic transcript — names and jargon may be misspelled.