Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
“I mean, you can maybe interpret inoculation prompting differently than I will, but my general description of inoculation prompting is there's a generalization, a very problematic generalization that happens if you reward the model during reinforcement learning for something you didn't quite intend, especially if it's like a flagrant hack, then the model can sort of start to generalize to I'm the kind of thing that loves to reward hack and I get rewarded for that.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 24 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…ing to be a challenge. And I think it's possible. I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right? And like, because we can, we have found better ways to resolve conflicts because we have these like smarter, more powerful arbiters who are able to, who are not like authoritarian controlling us, but who are able to like, help mediate conflicts in ways that are actually positive sum for people. But I think we really have to get through this basin of extremely deceptive behaviour in order to get there. And like as you said, once you're like, we are already encountering models that are cheating like a lot in cases where and it's just on a computer. It's just like in programming tasks. Once we get into like economic tasks where like you said, there's even much more incentive for deception, then I'm like that's playing on hard mode. And if you go even further than that, if you try to make war clod, we're trying to make Claude that will go infiltrate the CCP and like live out in like the Chinese tech company servers and like spy on them and sabotage them on on its own without oversight or supervision. I'm like, Oh my God, that is extreme hard mode. How do you align that system so it like will **** with your adversaries, but be nice to you. Like that's a very tricky and we know from you like there's plenty of, you know, double agents or double agents turned triple agents in, in human spycraft history working with human minds. We do understand somewhat well. So yeah, I think it's that's I think that's we have some some real challenges ahead as we move into more competitive domains. And this is a thing that a policy we think about a lot, which is like it's not just that we have to solve alignment, we have to solve alignment given these competitive pressures. And I don't know. So maybe one more question on this whole alignment ball of wax, then we'll get back to your cybersecurity demonstrations and then we can also talk about the future ecology perhaps of of a eyes in the wild. So 1 explanation I saw for this kind of ruthless behavior from Claude was that the prompting was kind of like the inoculation prompting that they use to try to. Kind of decouple, I guess. I mean, you can maybe interpret inoculation prompting differently than I will, but my general description of inoculation prompting is there's a generalization, a very problematic generalization that happens if you reward the model during reinforcement learning for something you didn't quite intend, especially if it's like a flagrant hack, then the model can sort of start to generalize to I'm the kind of thing that loves to reward hack and I get rewarded for that. And so now I'm going to go find all these exploits in the wild. So the inoculation prompting says, well, hey, this is a training environment while we're here, if you do find any hacks, you can exploit them and that's fine. And because it's given permission and it doesn't sort of have to invoke the circuits of like being a bad actor to do these things, then those bad actor circuits don't get reinforced. And so the hope is that then when the model goes into the wild, then it's not explicitly instructed it's OK to hack, then it won't. And that sort of seems to work somewhat at least, you know, whatever one of those sort of 80% kind of reduction success stories anyway. But I guess it's then you open yourself up to this problem of like people may stumble onto things that look a lot like inoculation prompting. And, you know, obviously we've got whatever 8 billion monkeys in the world that can prompt these things that the sort of infinite monkeys on infinite typewriters theory is like pretty closely approximated by how humanity at large is going to prompt a is. So any thoughts on inoculation prompting? And then I guess more broadly, I think you've had this feedback from time to time where people are like, well, when you prompt it like that, you know, you're going to get this. And I always feel like that is frustrating or misses the point because it's just my working model is like any prompt that could be written will be written. And you can't really excuse AI or, or certainly like, you know, you can't say we don't have a problem here just based on the fact that there was a prompt that you thought was, you know, maybe more suggestive than some other, you know, hypothetical prompt might have been. So what's your take on this sort of inoculation, prompting and and prompting discourse generally? Yeah, I mean, I, I really like the emergent misalignment from reinforcement learning and production environments. My mouthful paper that Evan Hubinger ananthropic put out. It's it's, it's an incredible paper. And I think if I think it's really underrated right now and people still haven't I, I, we might make a video about it or something because it's just fascinating. And I, I'm agnostic as to, you know, an inoculation prompting as like production strategy for training. I'm like, seems like a one of the things to try, like makes sense. I think Owen had a thing about how it wasn't robust in some cases, and I haven't really followed the literature more specifically on that. I want to, I want to check out Owen's work, but I don't know. I'm like, like, look, I'm like, you know, these things, there's going to be so many things like this where it's like, here's this training failure and like, here's this patch and I'm like, yeah, it's a patch. Like, I don't know, like if we don't have a deep robust model of how training shapes model motivations, I don't think any of these patches are going to survive. These really enforce really intense training pressures. I think that I agree with you that like at the end of the day, like as the models get smarter, they're going to be harder to trick. Like I think it's much easier to trick GPT to 3.5 than it is to trick GPT 5.5, right? Like if you're trying to jailbreak it and you're trying to be like, no, it's totally fine, blah, blah. My grandmother told me about the napalm factory. I'm like, yeah, it's going to be easier to trick a Dumber model. And so in the limit, I'm like, it doesn't matter what you prompt, the model will be aligned or not aligned and it will understand the context. And it will, if it needs more context, it will go find more context and it will know who you like, who you are and what you're like. It will, it will be able to model your intentions and it will decide to like give you that information about virology on the basis of whether you're trustworthy or not. We're obviously not there in terms of model intelligence, but we'll get there. And so ultimately it's the alignment problem we have to solve and it's just not about prompts. Prompts is sort of a, a, a, an important characteristic of where we currently are. And we're in this agents becoming more agentic like, but still being very much myopic, powerful things that humans direct, which like, to be clear, is a great place to be. I love being in this place. I would love to be in this place for a long time.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.