# From Models to Machines: Scaling Physical AI | Bayley, Ming-Chang, Travis, Amber

- Channel: [Ethereum Denver](https://streameth.org/ethereum-denver)
- Date: 2026-03-09
- Duration: 25:55
- Topics: ETHDenver, Crypto, Web3, Blockchain, Event, Conference, ETHDenver 2025, ETHDenver 2024, Bitcoin, Ethereum
- Watch: https://streameth.org/watch/yt-Yo8OUcO6274
- YouTube: https://www.youtube.com/watch?v=Yo8OUcO6274

## Description

🚀 Get Ready for ETHDenver 2026! 🚀

We're already hard at work preparing for next year's biggest Web3 event!

Keep your eyes peeled for more info on ETHDenver 2026—it’s going to be epic! 🌟

## Transcript

what I actually want. &gt;&gt; Hi everyone and thank you for joining us today at ETH Denver for our panel on physical AI. Um I am Amber. I will be moderating today. This is Owen and I'll start by letting our great selection of um panelists introduce themselves and what they are working on. &gt;&gt; All right. Thank you very much Amber. My name is Bailey. I'm the CEO of Prisma X. We're building the service layer for AI native robotics things like data collection operation teleoperation standards for robots that are actually useful in the real world. Uh we use blockchain to go coordinate all of these things run our incentives and store provenence and reputation. Uh my background is I've been in robotics since 2012 2011 and I've been somehow mining bitcoins even even earlier. My first bitcoin was a dollar but sadly I sold it and used it to buy a pizza. Uh so here we are now. No regrets. Uh happy to be here. Hi, I'm uh I'm Travis Good. I'm the CEO of Ambient. Uh you know, Ambient's goal is to make machine intelligence global shared utility infrastructure. We think that integrating with the physical world is a big part of that. Uh we're a useful proofof work network. Uh so we're actually using verified inference fine-tuning and pre-training on a single large language model that we run on all the nodes of our network uh to uh serve uh intelligence uh to the world. Uh you know my background uh is diverse. I have a PhD in IT with a specialization in ML. Uh I was in biotech my last career. ARC uh sold a drug discovery pipeline uh to SC Johnson. Uh prior to that I was the head of machine learning for uh Union Pacific Railroad where we did a lot of uh high-scale safety critical ML work. uh so that's me. &gt;&gt; Yeah. Hi everyone. My name is Ming and uh I'm a founder and CEO of APEL. So what we are doing is a agent to a an agent to agent communication and a behavior checking layer uh for for the agent agent to agent communication for business operation for example buy and sell or negotiation. So we'll connect two agents and then um detect any erroneous behavior and then inform human in the loop and then finally start uh certify the interactions. Yeah. So my background is in um in computer science. I got my PhD uh in AI as well and then uh previously I was a researcher at Google research uh and video research and some of my multiple works including some predecessor work for uh Google's V2 and V3. Yeah. and then uh training some like uh training and evaluation of uh some internal gymd5 models at Nvidia. &gt;&gt; Thank you. Um Bailey for the first question, what does physical AI mean to you and um mean to each of you and why does it matter right now? Starting with you, Bailey. &gt;&gt; Yeah. So I think physical AI is a very clear clear-cut concept. It means uh software for robots to make them more autonomous. Uh basically the idea is we have the hardware now. Robots have been around for years, but program them is really hard. You have to be an expert. You have to understand the machines and you have to go through this really laborious process in order to make the robots do anything. And if anything about the environment changes, you need to go change the code, right? That's really bad because it means you need a million dollars of engineering to be able to deploy robots. The idea is by leveraging the same statistical techniques used in things like language models, we can get the same level of general deployability in this robot software that we do in these language models. So instead of having to go rewrite the entire program to move the robot to a new location, you can simply prompt it with natural language or it automatically learns based on its environment. It's really important because it's what enables robotics to move from a multi-billion dollar industry to a multi-t trillion dollar industry. Robots are already worth a lot of money. They like every car you drive is built by a bunch of robots, right? But really that's all they do. They're trapped in these manufacturing lines. And being able to give this level of autonomy and a much more efficient user interface means that robots can go cover the other10 trillion dollars of manual labor in the world that we really really do do need solved. There's a labor shortage in a lot of countries. There's a lot of problems, a lot of dangerous situations. So having autonomy is what will make these robots be able to tackle everything from household tasks to grocery store stocking to more manufacturing tasks. I totally agree with everything uh that Bailey said and what I would add to it is that physical AI is a frontier that we need to conquer uh in order to really embody agents in the world and make them fully useful in the way that we want them to be. And it's also the convergence of sort of three different areas that are each very interesting. So you know one of them relates to large scale data acquisition uh for robotics applications. Another relates to large language models for uh planning and intelligence. And a final uh area is that of world models uh which express complex relationships among uh you know objects and entities within a world. And I think that physical AI is like right at the crux of all of these and is a requirement for us maybe to reach AGI as well. &gt;&gt; Yeah, I totally agree what the other two guys are saying but I would just uh provide a very uh let's say simplified mental model at least for myself. So I see physical AI I mean at least optimal ultimate physical as so what human can do you mean for for everything real time and I I I like to use the analogy uh in agent because agent is more closer to to realization. So if you think like agent is what human can do in front of a laptop in a virtual world. So physical ultimate physical is like what human can do every uh in the real world real time. Thank you. Um, next question is why we've seen massive progress in language and image models. Why has physical AI been so much harder to scale right now? &gt;&gt; So, I think the sort of obvious answer to that is there isn't enough data. And that's sort of pretty true. It's a pretty subtle question. What actually is is that it's difficult to sort of control the pre-training data. There isn't an obvious default pre-training data set like for language models because you can't download a an internet of the real world, right? there's no digital representation of the real world yet. So, uh the lack of pre-training data limits a lot of the experiments we can do. The models are a lot smaller, train on much smaller data sets. I think there's also some practical considerations. I don't think you could make a trillion parameter physical AI model really work practically, you would need a few hundred kilowatts of computers to run the model in real time. And that's a problem, right? Even if the computer electricity were free, uh a quarter megawatt being pumped into your house would make you very uncomfortable. While the robots would work very well, you probably wouldn't be able to live with that building anymore. So, I think there's some engineering considerations surrounding like power use and architecture. But I think Travis brought up a really good point. I think there's also a bit of a architectural limitation in how we think about the models right now. Like what are the representations for the physical world? We've sort of defaulted on the correct representation is a compressed video. And I don't necessarily think that's the right answer. In particular, you can't reason in image space. And if you walk around, if you're like trying to complete a task, you will think about what to do before doing it. You don't start waving your arms around seeing mistakes happen and then like undoing those mistakes in order until you arrive at the right conclusion. Right? So I think some sort of better joint representation for the real world, the language, the reasoning, the logic is needed and I think that's a open area of research which I really hope the the big AI shops start focusing on in the in the next year. &gt;&gt; Yeah. And I'm exactly on the joint representation page. And probably the thing that I'll add is that, you know, we're at a moment where I I think it could go two different ways. And you're I'd really like to hear your opinion on this, Bailey, but you know, you're seeing uh companies like OpenAI start to produce devices that have cameras in them uh that you'll probably be released in the near future. We've seen some sort of failed prototypes of these types of devices that will really be able to capture a lot of video information and probably kinematic information as well. And I think the question that we have uh in the world is can we create some open devices uh that will capture an open data set uh that's similar to the corpus that we have uh on the worldwide web or are these going to be entirely proprietary data sets that are collected by proprietary uh devices and I I think that that is a huge bottleneck as Bailey says right now uh the data and it's really the acquisition of that data that may determine uh the future of the the marketplace as well. &gt;&gt; Yeah, I I guess I guess my thoughts on that are there's a there there's like a really strong argument that large-scale video pre-training should be the prior for all physically related models because you can capture a lot of information. For for some unfortunate reason, most of the guys who like have money and are spending it to train physical models aren't really doing the video pre-training approach right now. They take an opposite approach where they start with these really small data sets collected by remote controlling robots and then use that to start building their models. And I think that's just a commercial consideration. I really believe that data set needs to be open. Like the common crawl is open and it's critical to all AI models. The model trained on these video data sets will be even more important than the text models because they will be displacing 10 trillion dollars of human labor. Like 40% of the world is going to have to find something else to do once these models exist. And I don't think it's fair that a large Silicon Valley technology company which basically has no real understanding of how real life works goes and controls that. So I think there's a really strong argument that something decentralized something with more control or the ordinary contributors should be the one that builds that large data set. Yeah, I think you guys uh say all I wanted to say but I I think I I agree that the ar architecture uh of trending the the physical AI model is a it's a bottleneck one of other than the data set because so if you read follow literature so every day I see different opinions on like a trend how to train the uh physical amount is it world model is it like a visual language action model is it VA plus RL so there's no like definitive or clear technical path to how to train the like a chat GPT like uh physical model. So we we haven't reached there yet. But I think it's a actually a good thing because um it opens up a lot of opportunity for startups like uh like bailies. So so because when there's no like a technical clarity, everyone has a chance. So you don't need to compete with the big labs for resources. And I think other than the what uh what we we were just talking about I think uh the safety and ethical concern is also another reason. So I think I think uh because there's no recourse if uh a physical AI make a mistake. If phys if a robot punch me in the face I I will probably just die. So yeah so there's no like um you know Yeah. So I think that that that's other two reason uh that physical physical AI hasn't scaled uh yet. &gt;&gt; Yeah. Right. If if chat GPD gets your homework wrong, it's like somewhat annoying. If the robot slaps you in the face, that's probably more than just a minor annoyance. &gt;&gt; Okay. What do you think needs to happen in the next 36 months or so for physical AI to have that chat GBT moment? &gt;&gt; Let's start the answers in reverse order this time. &gt;&gt; Sure, that's going. &gt;&gt; Yeah. So, I think um first I we I think we need a really really good like a benchmark. So like in the I think we can just follow the uh how LLM or visual language model like chatgypt like progress because so so those like um those like models thrive on the benchmarks. So so for example this morning like Google just released like Gemini 3.1 or something. So it like kills all the on all the benchmarks. So we will we need to see like a really really good benchmark that shows that uh so the benchmark needs to be similar to what uh what human day-to-day like life life looks like and then so people would actually believe that okay it's actually uh useful in the in the in our household and then uh once we establish load like a benchmark that shows that that it works and then um that will lay a really good foundation for the physical AI and then we'll start uh improving on those benchmarks. based so those benchmark will give us some indicators of of how models should should uh improve on different areas. So I think those two are u what I thought &gt;&gt; so I think we need a killer demo and a killer data collection device and I'm going to break those down a little bit. So for the killer demo I think that a problem that decentralized AI has and maybe decentralized robotics has is people simply don't believe that it works. uh there have been a lot of failures uh that you know we could go into a long tangent about why that's happened. I think they're probably primarily related to the economic incentivization scheme. You know ambient is specifically focused on proof of work there but there have been a lot of failures in decentralized AI and I don't think that people really believe first of all that decentralized AI could power like real physical infrastructure in real time. So I think a demo needs to incorporate uh first of all that uh and then it needs to have a real-time robotics element uh with an open-source uh software stack uh and it needs to be operating in some like very public setting uh where people can see the intersection of all these technologies producing useful work for humankind. And so that's the demo that I think that needs to happen. I don't know what exactly the setting is. I kind of have a spec of the technical architecture that would be required. Um, and then I think that the other thing that we need is an open-source uh data collection device uh that lets people earn cryptoeconomic incentives like a very lowcost version of some of the things that are going to be produced by the closed labs and that's just to give us a chance to produce our own common crawl uh so that we don't all get sidelined. I 100% agree with that and if there are any deepin folks in the audience, someone really needs to go make this open source camera data collection widget. I don't know why it doesn't exist yet. It's not that hard to make just no no one's been working on it. But like sort of summarizing the two people two people before me, I think what we really need is to see more business show up on these show up in physical AI, right? You have these huge labs. They have raised $3 billion, but by and large you don't see them do anything other than produce videos. Like at this point, 1x Robotics is basically a V film production company. No one's ever seen the figure robot in real life. Boston Dynamics is mostly known for its cool YouTube channel and secondarily for the things that Hyundai claims the robot will do that it probably won't actually do. Uh but just this idea of actually taking the pieces and technology we have now and trying to ship something which you as a consumer even for a large amount of money might actually want is something that I think really needs to happen. Fortunately, we see a lot of trends here. We're we we as a company are really sort of pushing in this direction. We started as a data company and we've been like trending more and more towards the rest of robotics because it turns out that once you start learning the nuances of data and what data sells, you also learn about what robots sell, what applications are useful and sooner or later you think to yourself, maybe I can work with a robot, other robotics companies, hardware manufacturers to build robots that can ship in the real world. And I think that goes back to Travis's idea of a killer demo, right? needs to be something where people look at the robot and they're like, "Wow, this is useful. I could see myself using this product, which is not something that any of the robotics manufacturers have delivered today. We've seen a hundred videos of folding shirts. If an alien visited the Earth right now, they would think that the greatest problem to mankind success is the fact that we have to not even do our laundry, but fold it afterwards. My god, this this this species worships neatly folded cloth, and they must fold the cloth neatly every day, and they've spent billions of dollars on folding the cloth. This is all kind of getting tiresome and people are just getting annoyed. They're like full shirt folding, dishwashing, back flipping. Someone needs to bring things together and like start building applications which are compelling and I think it's going to happen. I think it's going to happen in the next 18 months. It might come from us. It might come from someone else but overall it's just a great thing for the industry. &gt;&gt; Thank you. Um let's start this question the next question with Travis. What role do you see decentralization playing in scaling physical AI? Well, I I think you know, Bailey talked about the hard constraint that you can't put like a trillion parameter model in an embodied device. It just doesn't work. And so then you're needing to rely on some cloud AI to have high intelligence. Like if you're wanting to converse with a robot intelligently, have it make complex plans, have it respect your schedule, integrate with other things, uh understand the constraints of your environment and the context of that environment, you really need uh high intelligence. And that high intelligence is currently uh provided by companies that are essentially data black holes that are trying to use your data against you. Uh you know, we were joking before the panel about uh you know, the idea of having like Sam Alman in your bathroom. Uh it feels pretty dystopian uh to have these companies with these terrible privacy policies uh who are trying to obsolete humans uh be so directly integrated into your lives. And so, you know, I think that we need uh high-scale uh performant um decentralized machine intelligence that's credibly neutral that respects privacy uh so that uh these robots can have all the same features and capabilities as whatever the closed source uh versions are um but without creating a doom loop uh for humanity. &gt;&gt; Yeah, I don't so I 100% agree. I think I don't think I don't think it's going to be possible to go ship robot like cloud inferenced robotics models on the business model that a company like open AAI has right they have a ton of overhead they need to pay the research team have margins and they need to generate enough margins to support their share value which requires a lot of money and all of that just like for every cent of depreciated GPU price you might pay for inference open AAI is making something like a dollar and that just doesn't work when you need that inference running 24/7. Uh, so outside of the privacy concerns, which are very obvious, like Sam Alman is not getting anywhere near my bathroom no matter how much he tries, I think the just the overall cost efficiency of being able to push things closer to the edge where it looks like you you own the robot and you own the GPUs but not necessarily have to like physically deal with the servers in your house. I think that makes a lot of sense. So bas basically instead of the inference vendor right right now open AAI pay charges you to pay a cloud provider who then makes 10x the value of the GPUs over the lifetime of the GPUs. So you're paying 30 to 50x the actual cost. Uh by pushing it to the edge you can get those margins down to something way less. Right? Like ETH miners back when Ethereum was proof of work were happy to like we get 2x the price of the GPU over the lifetime of the GPU. They had tight they had really tight businesses. They were able to sustain those cost models. Like the GPUs were literally under tarps. They would rust away over the course of the three years. But that was fine. Like you knew ETH was going to lose all its value in three years anyway. So as long as the GPUs rusted away after after the crash, life was good. I think we we we need to see something like this. Like argue one one argument is okay, you ship the GPU with the robot, but it's like really annoying to deal with a large GPU server. It makes your house hot. It's very heavy. It costs a lot of money. So being able to push this out to someone else but still sort of maintaining this uh uh like closely knit connection between the robot and the GPU I think is really important and I think that can only happen with a lean decentralized style provider. &gt;&gt; Yeah, I think we know that the the chips are getting faster and faster and better and better and then probably could be cheaper in the future as the technology advance. So, I think um it's gonna for for me the natural way to think of physical AI is going to be let's say a bunch of robots interacting in a peer-to-peer fashion in instead of a client and service fashion. So, in a peer-to-peer fashion uh interactions, it only makes sense to have a decentralized um brands or something. Yeah, it's very scary to have a one like a vendor for example if it collapse then every every every uh robot collapse. So yeah, so decentralizing I uh decentralization I think is a necessary steps. I mean one one flywheel is that the the hardware size keep advancing and the other side is that it only makes that it's only natural to have a peer-to-peer interactions in the physical space. Yeah, I guess the only thing I would add to that is um you know, you do something like proof of work, right? It's self-healing. Uh you know, China kind of cancelled Bitcoin at one point and the network kept on chugging. Uh and the reason is that the economic incentives were in place uh such that uh miners emerged uh to help out and you need that availability and reliability uh profile for high intelligence supporting devices that are performing critical functions in your household. and uh to for economic good. Uh so I I think that's where decentralization can really uh be helpful. It might not always be the fastest thing, but it can certainly be the most resilient thing if you design it correctly. &gt;&gt; Thank you. At least Owen's head didn't explode. &gt;&gt; Yeah, I think you ran out of batteries. &gt;&gt; Okay. Um Bailey, can we start with um what's the most common misconception the public or investors have right now about physical AI that you've seen? I think there's a big public misconception that physical AI is about making robots do back flips and other forms of highly dynamic motion and that's largely it's like a weird story that back in 2017 18 when the modern robots first getting got started the benchmark was how high your robot could jump cuz it was really hard to do that at the time and it was mostly academic discourse right so once those papers were published all the hardware manufacturers picked up and for some reason or the other decided to uh use like the the old academic benchmark marks became the new commercial benchmarks which resulted in a series of humanoids with like everinccreasing abilities to leap in the air. Like I think this year we finally can safely say that humanoids have exceeded human performance. Like uh expert human martial artists pro would have a hard time matching what the machines can do this year. But even so this just it's not useful and that's not the right metric of performance. The right metric of performance is can the robot do things in the real world. So I once again I go back to what I was saying earlier. I'm hoping that this year, next year will be really the years where people start focusing on real use cases and real businesses because I think we have enough technology building blocks to start building those businesses. Now, &gt;&gt; Amber, could you repeat the last part of the question? Sorry, I lost the &gt;&gt; What's the biggest like most common misconception people have right now? &gt;&gt; Yeah. Well, I think that to some degree it might be uh people don't understand how one field can help another field entirely or they don't fully appreciate that. So I think it's pretty commonly understood that robots can collect a lot of data and that's going to be useful for LLMs. But uh on the other side of it uh LLMs might be able to develop like new softbody physics uh which could be extremely helpful in the robotics domains. And so I would love to see a lot more uh cross-pollination of ideas and focused research programs uh related to uh each area advancing the state-of-the-art uh in the other because I think they are quite complimentary. Uh so that's that's the biggest thing I see. People often just talk about these in silos. &gt;&gt; What about you mean &gt;&gt; same question? &gt;&gt; Okay. So I think um I don't I don't have a better I think the the two guys to my right has has very great um suggestions and uh I think that uh for for me not maybe not the misconception but I think just the expectation I I guess I think physical AI going into household might take a bit longer than the than we expect because um I mean some challenges that we talked about in the previous questions and then also um ethical legal concerns. For example, I guess like a physical to to have physical AI uh going into our household, it it should go through like a rigorous uh testing like a drugs FDA approval, all that kind of stuff. So, it might take longer than we expect it to to to land in in real life. &gt;&gt; Thank you. Thank you. We're about to wrap up here. I know you guys have lots of questions to keep for after the panel, but since we're at ETH Denver and there's a bunch of um builders here, could you all like each of you quickly just say in the inter intersection of AI and crypto, what should builders be focused on building on right now? &gt;&gt; You want to go first? &gt;&gt; 10 seconds. I think scalable highly useful services that have product market fit and it can address uh the needs of today and thus show a path towards the needs of tomorrow. &gt;&gt; Really &gt;&gt; I think I think just uh learn the fundamentals and then think first principle what is natural and then what is uh absolutely needed to happen. Yeah, &gt;&gt; basically the same as Travis but I'm going to say useful and scalable instead of scalable and useful. I think useful is more little more important than scalable because a lot of people talk about scale but they often scale things which are no one uses. Awesome. Thank you guys so much and um let's hang out after the talk. Thank you.
