Evidence receipt / evaluation
Published · transcript-backedJeffrey Ladish: evaluation
24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
“And the scale up from O1 to O3 is incredible. And I don't know if O3 was even out yet, but like in that time when we had moved from just like a pre training regime where it was just like throwing a bunch of human data to the point where no, no, we can actually train these models by they can do trial and error on their own.”
Source trail
Everything needed to verify it.
- Speaker
- Jeffrey Ladish
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 24 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…This has been a long time coming. We've met a few times at different events over the years and I cross posted an episode that you did on another podcast some time ago and I'm glad to finally be doing one of these live. So it should be a very interesting conversation because you are right in the thick of it right now at the heart of where AI capabilities are going vertical and the consequences are going from theoretical to practical concern even for obscure folks like me on a day-to-day basis in a pretty compressed time frame. So I'm going to be interested to hear both in the weeds, details about the research that you've been recently doing and the observations that you guys have made at Palisade. And then really also looking forward to broadened out conversation on what can I do about this, if anything to protect myself and what does it mean as we go forward into the very foggy AI future? Yeah, well. Well, Nathan, I remember a year and a half ago, my team and I went to DC and the thing we're doing in DC, we're briefing a lot of folks in Congress in the admin, a lot of different staff. And we had a presentation that was like, hey, autonomous cyber agents are coming. Like AI agents that can hack pretty autonomously at scale are on the way. And the reason we know this is because Open AI just released a model called O, one that has been trained via reinforcement learning on actual programming problems. And the scale up from O1 to O3 is incredible. And I don't know if O3 was even out yet, but like in that time when we had moved from just like a pre training regime where it was just like throwing a bunch of human data to the point where no, no, we can actually train these models by they can do trial and error on their own. They can do exploration on their own. We have to build reinforcement learning environments for them. But once we have these, they're just off to the races. And we're seeing this incredible performance in programming. And now a year and a half later, we have Mythos, we have GBT 5.5. Even the open weight models are getting very good at hacking, as some of our work has been showing. And anyway, here we are. It's been an interesting couple years. Yeah. Well, let's start with a couple of the pieces that you've put out that I would broadly file under the scary demos category. I don't know if that's how you think about them or how you think of the theory of change on some of these projects. But the two that I wanted to cover that are most recent are shutdown resistance in robots and then self replication of agents. I think these are like in some sense, relatively simple and just striking examples of possibility, but it is probably worth getting into the details and and really making sure that that people are well grounded on exactly what is happening and you know how they should be interpreting some of the findings. So shut down resistance in robots. I'll give you the the quick my understanding and then you can expand on it. So basically you guys take one of these little doggy robots, so cute, and put an LLM in charge of it. So I'm imagining your classic agentic loop where the LLM is getting the state of the robot and what its camera is seeing, and then running its reasoning loop and deciding what to do. And then you put a little wrinkle into its environment, which is a big red button that is labeled dog shut down. And then strange things start to happen, the LLM seemingly identifying as the dog somehow understanding that if I, if the dog is shut down, I don't get to play anymore or I, and I can't accomplish whatever goals I might have, which is something maybe we should dig in to a little bit as well. What, what goals did you give them? What goals do they have? Goals. Anyway, the LLM takes action to prevent that button from working, like rewriting its own source code to intercept those calls or disable those shut down calls. This I think you know, should alarm people at least a little bit. But tell me more.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.