New Ethereum talks, every Monday. The week's conference uploads by event, in your inbox.

Loading player…

Running Wargames to Prepare Protocol Teams for Incident Response | Devcon SEA

DevconTue, Oct 7, 2025, 12:00 AM

SEAL (Security Alliance) Wargames: cybersecurity exercises designed to enhance Web3 protocol resilience. We'll share experiences from running these with major Ethereum protocols, covering: -Exercise structure: OSINT, tabletops, and live simulations on forked networks -Scenario designs and common vulnerabilities -Infrastructure and open-source tooling -Key learnings and best practices -Scaling strategies and the importance of regular security drills in the evolving Web3 landscape Speaker(s): Isaac Patka, Kelsie Nabben Skill level: Intermediate Track: Security Keywords: Coordination, Security, incident, response Follow us: https://twitter.com/efdevcon, https://twitter.com/ethereum, https://warpcast.com/devcon Learn more about devcon: https://www.devcon.org/ Learn more about ethereum: https://ethereum.org/ Visit the https://archive.devcon.org/ to gain access to the entire library of Devcon talks with the ease of filtering, playlists, personalized suggestions, decentralized access on Swarm, IPFS and more. Devcon is the Ethereum conference for developers, researchers, thinkers, and makers. Devcon SEA was held in Bangkok, Thailand on Nov 12 - Nov 15, 2024. Devcon is organized and presented by the Ethereum Foundation. To find out more, please visit https://ethereum.foundation/

Transcript

[Music] [Music] and I also lead the uh War Games Initiative for the security Alliance um I'll give a brief introduction to just what the security Alliance is and what the various initiatives are um because I think it's actually pretty relevant also after the first talk um about ens and data sharing and so the top initiative here is actually called the seal ISAC which is not me I'm seal Isaac this is the information uh uh security something something which basically this is like a uh a shared database between many different companies where they can share data on uh like dprk threat actors or fishing data or bad domains um there's seal war games which is the initiative that I lead where uh we prepare protocol teams for incident response through both tabletop simulations and then live simulations on forked environments um seal also operates seal 911 the emergency hotline if your protocol is under attack or if you have discovered a vulnerability a protocol and need to get in touch with um a security expert as fast as possible um The Safe Harbor uh agreement um which is legal protection if you're a wh hat hacker and you're acting benevolently in an urgent situation to uh rescue funds from a protocol think uh Nomad Bridge hack a few years ago if you were trying to save uh money but maybe you were worried because you thought that you'd be sued for hacking during uh because you it be impossible to distinguish the white hat Safe Harbor agreement exists to protect white hat uh hackers uh in those situations and a legal defense Fund in case they sue you anyway um so today what we'll talk about primarily is war games so what are these war games and why do we do them um I'll have some links to resources on how you can work with us uh there's a lot of open source toolkits and we also run these um as a as like a public good uh throughout the space um we'll be talking through some best practices that I've learned from working with some of the major protocols in the ecosystem um and then just briefly show you what this new toolkit is that we released so you can run these yourself uh if you would like so what are war games these are cyber secur exercises for web 3 um I've heard uh traditional cyber security folks call these something like a purple team exercise where we're not exactly red team blue team but we're basically just trying to stress test these teams and understand how they behave under pressure and so we want to help teams practice for these um high pressure situations before they actually happen um we want to be able to like test both the technical side of their team and their social resilience and we're not just trying to help like cause them to panic but we want them to practice emergency response um a lot of the uh a lot of the reason why we started doing this um was a few years ago um so the security lines was started by Sam sun and he was pulled into so many War rooms where it seemed uh like I guess in uh uh from what I've heard uh teams just completely melting down and not really being exactly uh prepared for what would happen so what could we actually do to help put them in that situation before that inevitably happens um so War Games take place through three phases um I do a lot of uh uh intelligence gathering which really just means reading every document that the protocols ever published understand their contracts understand how it works uh we do a tabletop exercise where we talk through various scenarios with them um and understand how they would have detected them what they would do who's in charge um and how would they have fixed something and then what I think is the most fun is we do a live simulation so we set up a forked environment we run all of their contracts we run all of their uis we run all of their monitoring um and then they actually have to coord in real time to uh uh defend and respond as if it was a real incident and we tend to find a lot of things that they can improve through that process and it also just helps a team build trust in each other so that they know when an incident goes down for real I can trust my team uh uh to be there so a big reason why we do this is because more and more hacks are due to operational failures not smart contract bugs that's a really I think good sign that we're making fewer and fewer dumb mistakes on the smart contract side how however as all of the stuff that we're building is becoming such a critical component of Global Financial infrastructure the stakes are higher to actually operationally be strong as well and so whether that's uh whether that's things like supply chain attacks or uh or social engineering uh it's much harder now than just like not uh including dumb bugs in your smart contracts so these are cross functional exercises we're not just working with a development team they're of course a core part of it but we're also tending to work with their with work with their Auditors if they're a big protocol that has a guardian multisig that's in charge of pausing the protocol in case of emergency we of course work with them um Communications team is actually essential so that we know like what would they be communicating out to the public and when during an incident and legal so that they can say okay if we were to take this action perhap to perhaps to protect the users of our protocol can we do that should we do that and what should we be saying the conversations between Communications legal and devs should really be that are figured out in advance so that during the incident you're not thinking like do we tweet out that we're under attack right now or do we say something cryptic or what do we say like these are things that you can just have in a Playbook so that you don't have to think about it uh in the moment we've worked with a number of products uh of protocols uh so far uh but the purpose of this slide is also just to show you like why teams are working with us it's because like teams are growing uh these are Global teams and like it's really hard to like prepare for this type of uh this type of thing um there's many ways to engage with us if you're a protocol that wants to work with a security Alliance and have a drill performed with you you can uh there's a form on the on the website and there's also open source uh uh resources for you to be able to run these yourself so uh this is a quote that I had a really hard time attributing uh the internet said it was either like a Greek poet or the Navy Seals but I think that it really like uh stands to to show like why we do this and it's uh that Under Pressure we don't rise to the occasion we sink to our level of training uh yeah I think that that speaks uh to why we do this so each one of these builds is quite different um a lot of protocols uh a lot of protocols have like various sorts of like custom infrastructure and monitoring so every time we do this we end up having to like build a lot of things and we've tried to combine all of those tools and things that we've learned on how to run these simulations um into this uh drill template that you can see here it's on the security Alliance GitHub repository but what are the things that we do to make the environment feel realistic so we of course have to run a network Fork we're usually either doing that with Anvil or with uh tenderly we have to run a block Explorer so we're either running block Scout um or using a virtual test net um we're running monitoring so it's essential to mirror the same monitoring that you would have on Main net so if you're using a product like hyper native hexag or Forida you should be we should be setting all of that same stuff up against our forked environment we write Bots to behave like users so if it's a lending protocol we have a lot of noisy transactions of users depositing borrowing Lending spping um doing everything that would exist on mainnet so that the only TR so it's not like the team only sees like two transactions and they're an exploit and then it's too easy we try to make it a little bit harder um we have an exploit which uh an exploit bot which ramps up intensity over time and then if there's any sort of miscellaneous things that they need to pull in we try to provide that um as well so this takes us usually a couple weeks to build for each protocol but we're trying to get faster and faster each time and part of that is uh uh thanks to this uh open source template that we're that we're working on so what are some of the scenarios that we've uh that we've worked on so uh as you've seen we've worked with various like lending protocols and decentralized exchanges and l2s and so the things that we're work that we like try simulate are not necessarily uh like bugs as in like a smart contract bug because we're not Auditors we're not trying to just find like a mistake in your code um and in fact like if we do that's like that's like not what we're that would not be the goal we're trying to figure out what are the things that can happen in if the protocol works even how you designed it so what if some sort of external uh dependency fails what if the way that you're importing Oracle data is weird or what if like some token uh some token contract gets upgraded and starts behaving differently so what are the things that are outside of your protocols control that can uh that can change what if your upgrade goes poorly how would you detect that um what if there's something strange happening with oracles and also uh one thing that I think was actually pretty fun is in some cases we did like a a malicious governance proposal uh C f with some protocols where we just like kind of flooded them with a bunch of governance proposals that looked like normal funding requests for events or sponsorships or just uh parameter updates but some of them also included like malicious data um like how tornado got uh hacked or I think a number of protocols how how they got hacked just by having like sneaky additional call data um that uh that ended up having wide uh bad implications uh I've pulled out a few best practices of what I've learned from working with with uh working with these teams that I just want to go over briefly and so one is just maintaining multiple communication channels to reach your teams uh one thing I like to tell teams is like imagine uh like during a tabletop like imagine this scenario happens but now imagine that this happened and telegram is down on the same day that this happened can you actually reach the Guardians do you have another way to reach them can you reach them on slack Discord can you page them uh sometimes we find that like when we work with a team like a a paging system that they think is set up is just deprecated and hasn't been working in a while and they can't actually reach the teams that they want to or they reach the team uh and they don't have their Ledger with them and so like okay cool we got in touch with you but you can't actually sign the pause transaction because you're camping like these are things that you want to test um also risk isolation this is a we learned a lot I think from the uh from how the UR team um prepares and isolates risk between the different strategies and so there's there's a a big lesson in like minimizing the amount of contagion that can spread uh from like one strategy to another and also the diligence process before any strategy was employed there was always an emergency response of like with this change these are the things that can go wrong and these are the steps you would take to recover funds before they become uh unrecoverable so like Risk isolation with urine uh similarly like Asset Risk isolation with a we found was very interesting how you can see how there's all these different configuration Flags based on how potentially risky a different collateral asset is is they have different properties as far as how they can be uh how they can be used in the protocol so I think that these are all really interesting lessons from different like larger dii protocols about like how you minimize potential contagion and and and things from uh things from spreading uh finally uh or additionally um incident post-mortems are really important so uh this is another call out to like the urine GitHub repository where you can see that every time there's been something that has happened you can uh see what they what the team did how they found out and then also how they'll be improving in the future to uh to avoid that sort of issue and monitoring uh mirroring monitoring I think is quite important like a lot of teams only will run like their critical monitoring against their main net smart contracts um and that's handy but uh the things that are actually going to trigger to maybe like pause your smart contract in the event of emergency maybe it will never happen or maybe it will happen like once a year and so how can you really be sure that it's going to be there the moment of and a lot of the these these tools are great and I'm sure that they have some internal simulation capabilities but if you can also run your monitoring against these like simulated networks so that you can make sure um that it's all functioning the way that you think it's functioning um I think that's that's pretty key so in the final few minutes um before I'll leave a few minutes uh we'll have a few minutes for questions I just wanted to show you the the open source repository uh that I la that I released a couple weeks ago with all the tools that I use to run these for protocols um so it includes like a Foundry and hard hat setup for like validating different like war game scenarios um on a local Fork it contains a bunch of configuration scripts for setting up virtual test Nets with block explorers to run these uh templates for the tabletop exercise uh monitoring and Bot Services which we forked from um an optimism repository and also like a stack for setting up like a Prometheus grafana instance so like this is all that I use when I run a when I run a war game so this should be all uh that you would need to potentially use if you want to run one of these with your own uh with your own protocols so this is just a a brief look at like what the tabletop script looks like and so we'll be on a call usually uh it's like an hour hour and a half with like the whole with the whole protocol team or with like if it's a very large protocol maybe it's a subset of the protocol uh a subset of the team um but always cross functional and we just talk through like in this exercise we're going to um discuss various scenarios that you'll need to detect and respond to issues what we want to like uh understand is typically like uh do you know what the key dependency of your protocol is and you know what would happen if it fails is there sufficient monitoring in infra to detect and respond to these failures um do you understand who has admin capabilities within your within your protocol um are there backups in place so these are the this is almost an additional information gathering step for us um so that we can design the best uh live scenario and then the questions that we ask in each one of these scenarios is always um like how did you discover this issue so if we're looking at maybe an an L2 and somebody submits some sort of fraudulent uh withdrawal proof how did you find out did you wait until somebody rate like uh reached out on Twitter or Discord and said hey it looks like you know XLT is under attack right now or did you have some internal system that would have alerted you within seconds and so first how did you find out do you know who to gather do like does somebody become an Incident Commander here do you have like an on call team or is it like whichever Engineers awake somewhere in the world is the one that has to to deal with this like is that uh how does that work do you have a response plan for this if not why not is it just too rare um and then who took action and how and so all of these questions are how we can understand like how well positioned a protocol is um to uh actually uh to respond uh and then for the live drill um the setup steps this is always the part that takes me the longest but we create the fork run the Bots brief the team I'll send them some sort of like uh atic message of like you know you're now in like a very you're now in a very realistic simulation where this really bad thing's about to happen you have to treat it as if like a everything was uh as if it's all real life um but you typically don't need to encourage that like they'll go into stressful mode and uh everybody wants to show that their internal tooling really works really well and so um uh and so yeah teams we found always take it very uh very seriously um I always love love to observe kind of the especially just the conversations between like Dev and comms where it's like Communications prepar is like a statement developers are like no we can't say that at all or developers want to say we're pausing the protocol and Communications is like no we can't say that like so that's always really interesting uh to to observe how that works uh we have um uh this is just a view of like a tenderly virtual test net so I found that this is like I used to set up like an anvil fork and a block Scout uh drop uh and a block Scout Explorer um that would always take me like two days um So lately I've been uh using these tools but you can run these Forks using any any tools but this is like what it looks like when I set up a a fork of main net um and then this is like my grafana setup where um I'll have like an in just a basic invariant monitoring thing where um we're checking for certain conditions like you know more tokens were withdrawn from an escro contract than we're actually locked up for a user and why would that happen and and how would we fix it uh and then the uh we the team actually has to submit the transactions on the test net um we try not to we try to uh do a couple things to keep safe like change the chain ID of the test net so that the um so that the actual like response transactions couldn't actually be replayed on Main net because we wouldn't want to be able to like take a pause transaction from our test net and then actually pause a major protocol on Main net and so we change the chain ID um and then we also try not to uh make them access like keys for safes and stuff if they really don't need to and so we'll kind of spoof the process of uh of doing like a multi transaction just in case like there's all sorts of security procedures around accessing these Keys um yeah so we try to try to keep team safe doing that uh here's a link to a few resources the first one is uh actually related to an event that was two days ago that we hosted uh here in in Bangkok but I left it there because we I think we'll be doing more of these like Live Events where you can compete with each other uh to try to white hat rescue funds from protocols uh the second is the GitHub repository uh we have the security Alliance website where you can uh sign up for a drill sponsored by the security Alliance um and you can also sign up for drills um on my company's website Shield 3 uh if you're on a an accelerated timeline uh or feel free to chat with me if you're just like hey I want to run war games for protocols uh I'm happy to share so thank you all right thank you very much Isaac appreciate it even though we started late you finished well within your time I really appreciate that and I'm happy to say that our questions are now working so let's turn our attention to the screen I'm going to read out the first question with the most votes is it possible to have a Seal Team Member be part of a protocol Guardians multii somehow y yeah um a lot of us are on some of these uh multisig or on security councils we actually talked about this recently like if seal itself should be like members of security councils we came to the conclusion I mean this might change in the future that like us individually can also be like on security councils uh but seal itself is not but like many many seal members are open to being like on these uh uh on guardian multisig and uh and stuff like that so feel yeah just find me and we can we can help all right thanks for the answer thanks for the question next question that got voted to the top do you share threat intelligence data on chain if no how or where do you share threat intelligence data yes so the security Alliance website should have a link to the cal- ISAC which is the information sharing tool um there's a there's a a large amount of data that actually is available publicly but then also if you're like a wallet company or an exchange where also um there's like even more sensitive data that can be shared in a more private way um so that's the best place to look for like uh for just data being shared across teams it can be it there's they're even starting to gather a lot of information just about every time like a a North Korean developer applies to a protocol they try to just like profile it and understand like how they're doing it who's getting in just so that we can like share data on like how teams are uh how teams are being like trying to be infiltrated thanks seal ISAC yeah thanks for the answer great all right let's go on we've got time for more questions does the drill occur unexpectedly I.E the team knows it'll be happening or but doesn't know when uh I've gone back and forth on this a few times initially it was like we would give them an 8 Hour window and then say like okay it's going to happen at some point in this 8 Hour window um but then I was just sitting there like stressed waiting for it to start so I always just started it right at the beginning of the window um so I don't know if like uh it might be fun to do that I think at at some point we're talking we're thinking about like can we do like more surprise drills um on these longer running test Nets where we like submit a faulty or like we submit like a a faulty proof to like an L2 rollup or something like that just out of nowhere um just to surprise a team um but yeah typically they know that it's going to happen on a certain date um but it I it still maintains a pretty good amount of realism right thank you very much surprise you've been compromised anyway uh let's move on to the next question I think this is an important question it's got quite a few votes as well what's the most common way teams fall over I think one of the most common things is like having your internal tooling just not be up toate um so it's very like maybe like one developer got on a was very inspired a year ago to make all of like these really great internal tools but then like the protocols changed a bit over the last year and the internal tools don't work so I think like maintaining the internal tooling that you need need for these scenarios is is very important and dedicating engineering resources to doing that so that like it's not like oh I really hope this tool Works um when you're actually in the incident excellent answer fantastic we got a few a little a little bit more time let's try and do one or two more any way to react to layer one protocol incidents uh yes so I think that we want we're thinking about some fun things to do on the layer one like maybe related to um like major protocol major like L1 protocol upgrades thinking about what could happen if there's major chain splits or Forks or bugs in a certain consensus client but not another one like what chaos could we simulate across the whole ecosystem if upgrades perhaps go really poorly or if there's bugs or supply chain attacks into major like uh validator code um of course like these are things that shouldn't and probably won't ever happen but it's fun to practice uh and I think worthwhile to practice uh the response excellent Isa you got one minute for the last question do you go through the worst case scenarios in war games for example you got hack money's lost what now do you go to contact Law contact with law enforcement uh typically I actually that's a really good question which I didn't touch on which is like we try not to just make these like an instant wrecked scenario because it's no fun if we just say protocol was upgraded there's a bug all the money's lost you can't do anything like that's no fun for anybody so we try to make it something that's like recoverable that they could actually respond to um but uh I think yeah actually on the last one um uh where we had like some money stolen they were like okay now what we would how would we actually try to like Chase down this money that we got that was that was uh was stolen and the answer is like that's something you can go to seal 911 for uh because they have really good like tracing people that can like say okay now we're going to see where the money's gone and try to recover it in other ways um and so uh yeah if you get hacked uh uh first fix the protocol but then yeah we can also help you uh track down the money excellent wow we finished all the questions in record time let's give Isaac a big round of applause well done thank you so much Isaac well

Automatic transcript — names and jargon may be misspelled.