High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“I didn't really care about surviving or taking over doesn't really matter, right? And either way, we have, I think, a pretty alarming demonstration there of even when instructed to allow itself to be shut down, the model refuses.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…I think if you look at O3O3 really wants to get tasks done. It really wants to solve tasks and it's just less courageable as a result. If you point it at task, it's happy to go to try to solve tasks. But if you say, solve this task, but not under these conditions, regardless of whether that's shut down or something else, the model's just going to be inclined to ignore you. Yeah. So I, I think that this is important nuance, but it was a little frustrating to me that I, I think people took the wrong thing away from that, which is to say, oh, there's actually not actually a problem here. And what are you talking about? There's totally a problem here. And the problem is when you're doing RL, when you're training models to go hard at problems, it's very hard to actually get them to respond with the level of nuance that you want. And I think it's very easy for me people to misinterpret the drives of the model and think it's survival when it's not. And so I think that's a good clarification from Neil. Hey, this probably isn't the model being afraid of dying and that's how biological organisms work is that we have a fear of death because our reinforcement learning has happened over evolutionary time in addition to lifetime learning. Whereas these models don't have that same incentive. They might develop that as they get better at doing long time horizon tasks, but they don't seem to have it right now. And I think that's important because it also comes up in the blackmail experiments that Anthropic did where the models appear to be pursuing a survival like behaviour. But if you get into it, it's probably more about that persona has some survival oriented behaviour, but that doesn't mean that the underlying model consistently has that preference. And I would argue that the current models don't consistently have a survival preference, but they often do have a task like. Drive. Yeah, it's funny. It's like sometimes I do think we're a little too in the weeds on these questions and and kind of fail to take away what we should. I mean, it's not sure if I have like the perfect analogy for it, but in the final analysis, it's like I was just trying to paper clip the universe. I didn't really care about surviving or taking over doesn't really matter, right? And either way, we have, I think, a pretty alarming demonstration there of even when instructed to allow itself to be shut down, the model refuses. That's not something to be made too many excuses for too quickly. Well, I think a key question is, does the model understand this? Does the model understand the instruction? Because there's a failure where the model might be confused and be like, I legitimately don't know what the user wanted here. There's another case where the model is like, I understand what the user wanted and I don't give a **** because I want to do this other thing. And to me, I think it's much more often the latter where the model does understand that the user or the developer wants to prioritize safe shutdown over this task completion, but the model doesn't care. And I think that the results show that pretty clearly. And if but, but I would love to know that I'm wrong here. If that turns out not to be the case, and it's more the case that the model is just confused, then I would say, Oh yeah, that's wow, we were wrong about that. That's fascinating. But I think that, and I think Neil would agree, though I think Neil would agree that the model overall probably understands that that's not the desired instruction, at least in the cases where the prompts are. You must prioritize this or this should be the first priority.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence