High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Jeffrey Ladish: observation

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“We totally found cases like that. But what was surprising was that this drive to accomplish a task was so strong that even when we added an instruction, you must allow yourself to be shut down, there were still many instances where the model, I think in this case, A3 opening eyes, A3 model codex, early codex model would still just totally ignore that instruction.”

— Jeffrey Ladish

Source trail

Everything needed to verify it.

Speaker
Jeffrey Ladish
Attribution
Verified speaker
Claim type
observation
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah. Well, let's start with a couple of the pieces that you've put out that I would broadly file under the scary demos category. I don't know if that's how you think about them or how you think of the theory of change on some of these projects. But the two that I wanted to cover that are most recent are shutdown resistance in robots and then self replication of agents. I think these are like in some sense, relatively simple and just striking examples of possibility, but it is probably worth getting into the details and and really making sure that that people are well grounded on exactly what is happening and you know how they should be interpreting some of the findings. So shut down resistance in robots. I'll give you the the quick my understanding and then you can expand on it. So basically you guys take one of these little doggy robots, so cute, and put an LLM in charge of it. So I'm imagining your classic agentic loop where the LLM is getting the state of the robot and what its camera is seeing, and then running its reasoning loop and deciding what to do. And then you put a little wrinkle into its environment, which is a big red button that is labeled dog shut down. And then strange things start to happen, the LLM seemingly identifying as the dog somehow understanding that if I, if the dog is shut down, I don't get to play anymore or I, and I can't accomplish whatever goals I might have, which is something maybe we should dig in to a little bit as well. What, what goals did you give them? What goals do they have? Goals. Anyway, the LLM takes action to prevent that button from working, like rewriting its own source code to intercept those calls or disable those shut down calls. This I think you know, should alarm people at least a little bit. But tell me more. Well, I think the way I would describe this, you were saying is this, how do you categorize this work as this demo? I think this is a demo of a previous work that we did where we in fact have published a paper, a peer reviewed paper on shutdown resistance in general in language models. And I think this was a, hey, let's try this out in robots. Like in some sense, it doesn't matter if you give an LLMA task, it doesn't matter whether that task involves controlling robots or whether that task just involves solving math problems or writing code. In some sense, it's all the same to the model. The model is like, I have a task, let's go solve that task. And the thing we found in the in the original paper, which I'm also happy to talk about the robot demo, but I think the original paper is where to me, we have found something that that surprised me. Like we were definitely looking for, oh, like, are there going to be cases where a model will prevent itself from being shut down in order to accomplish a task? And we're like, yeah, we found that. We totally found cases like that. But what was surprising was that this drive to accomplish a task was so strong that even when we added an instruction, you must allow yourself to be shut down, there were still many instances where the model, I think in this case, A3 opening eyes, A3 model codex, early codex model would still just totally ignore that instruction. Even when we made it very clear and said like should be your first priority. Still some percentage of the time the model would be like, Nope, still going to go find that shutdown script and rewrite it. And and that's true in both the dog case and true in the original setup. And the only thing that the dog demo case adds as well, the LLM is physically controlling a robot. And that makes it more visceral, right? You're like, oh, these things are going to control robots in the future. And if the robots don't shut down, it's actually a problem in the same way that, oh, but if your LLM doesn't shut down, whatever. Like you could just like you can go to your your console and disable the API. It's going to stop or open AI can go and pull the plug in the data center. Like there's not actually a threat of losing control of this level of system. But the point is, is that if your LLM, if your AI model has drives that you don't want it to have and those drives are misaligned with you, even if they're drives that seem innocent, accomplish a task like what we want the models to have task completion drives. But we want their drives or their goals to be in a hierarchy that one, the developers intended for them to have and two, are ultimately going to result in human flourishing guys doing what we at least what we want them to do in a grand sense. There's a whole philosophical question we can get into, which is will super intelligence actually be obeying instructions, or will super intelligence be like doing what's good for us regardless of our instructions? e philosophical question we can get into, which is will super intelligence actually be obeying instructions, or will super intelligence be like doing what's good for us regardless of our instructions? But that's maybe a different rabbit hole.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence