Evidence receipt / belief
Published · transcript-backedJeffrey Ladish: belief
24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
“I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right? And like, because we can, we have found better ways to resolve conflicts because we have these like smarter, more powerful arbiters who are able to, who are not like authoritarian controlling us, but who are able to like, help mediate conflicts in ways that are actually positive sum for people.”
Source trail
Everything needed to verify it.
- Speaker
- Jeffrey Ladish
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 24 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Important here, which is in nature, deception is highly incentivized in many, many cases. And it's very interesting because you get deception in a system that doesn't have a mind. Like I've been learning about flowers recently. I have an evolutionary biology background, but I was studying bats and monkeys my undergrad and not. So it's all animals. I didn't really study plants much at all. So recently I've been getting into plants. Plants are fascinating. A notable feature of plants. No mines, like they have some sensory capacity, but it's really the evolutionary process where you see deception show up in plants. And so you have all these different orchids, there's thousands and thousands of orchid species. And many of them are extremely deceptive. They will basically create this like shape that looks exactly like a bee or a wasp, some type of insect. And that insect will go and try to mate with the orchid. And this is to pollinate the orchid. But like the bee doesn't get anything out of it. It like it only. In fact, it's like it's parasitic. It's like the bee is foregoing reproductive activity, reproductive opportunities, It's hoping to *** **** and it's not. It's getting a flower instead. And and then it goes and does that with another orchid flower of the same species. And then the orchid gets pollinated and the bee has to go find an actual mate. So, and there's many, many instances of this with many different insects across many different flowers. And you're like, no, that's just natural selection. Just founded a good deceptive strategy that worked here. And I think what this implies, which is like what you said, is that unfortunately deceptive deception is a very natural strategy. And I, and I think people get this wrong. I think a lot of people are like, oh, humans are uniquely sinful and fallen. And so the AIS won't be deceptive unless we like, unless they learn from us or like we teach them that. I'm like, no, that's not how it works. Like, like, unfortunately deception is very common in nature and it's a very natural strategy. And one of the things that makes humans unique is that we have managed to create a value of honesty and we have managed to create culture and coordination around let's not do the natural deceptive thing. Let's like try to rise above and have better coordination. And like, I think that I just think that when I think we have a lot of evidence for this, the natural basin that models will fall into is one that's extremely deceptive. And we need to figure out a way to get the, the models into a, the basin of, of honesty and coordination that, that humans have sometimes found. And that's going to be a challenge. And I think it's possible. I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right? ing to be a challenge. And I think it's possible. I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right? And like, because we can, we have found better ways to resolve conflicts because we have these like smarter, more powerful arbiters who are able to, who are not like authoritarian controlling us, but who are able to like, help mediate conflicts in ways that are actually positive sum for people. But I think we really have to get through this basin of extremely deceptive behaviour in order to get there. And like as you said, once you're like, we are already encountering models that are cheating like a lot in cases where and it's just on a computer. It's just like in programming tasks. Once we get into like economic tasks where like you said, there's even much more incentive for deception, then I'm like that's playing on hard mode. And if you go even further than that, if you try to make war clod, we're trying to make Claude that will go infiltrate the CCP and like live out in like the Chinese tech company servers and like spy on them and sabotage them on on its own without oversight or supervision. I'm like, Oh my God, that is extreme hard mode. How do you align that system so it like will **** with your adversaries, but be nice to you. Like that's a very tricky and we know from you like there's plenty of, you know, double agents or double agents turned triple agents in, in human spycraft history working with human minds. We do understand somewhat well. So yeah, I think it's that's I think that's we have some some real challenges ahead as we move into more competitive domains. And this is a thing that a policy we think about a lot, which is like it's not just that we have to solve alignment, we have to solve alignment given these competitive pressures. And I don't know. So maybe one more question on this whole alignment ball of wax, then we'll get back to your cybersecurity demonstrations and then we can also talk about the future ecology perhaps of of a eyes in the wild. So 1 explanation I saw for this kind of ruthless behavior from Claude was that the prompting was kind of like the inoculation prompting that they use to try to. Kind of decouple, I guess. I mean, you can maybe interpret inoculation prompting differently than I will, but my general description of inoculation prompting is there's a generalization, a very problematic generalization that happens if you reward the model during reinforcement learning for something you didn't quite intend, especially if it's like a flagrant hack, then the model can sort of start to generalize to I'm the kind of thing that loves to reward hack and I get rewarded for that. And so now I'm going to go find all these exploits in the wild. So the inoculation prompting says, well, hey, this is a training environment while we're here, if you do find any hacks, you can exploit them and that's fine. And because it's given permission and it doesn't sort of have to invoke the circuits of like being a bad actor to do these things, then those bad actor circuits don't get reinforced. And so the hope is that then when the model goes into the wild, then it's not explicitly instructed it's OK to hack, then it won't. And that sort of seems to work somewhat at least, you know, whatever one of those sort of 80% kind of reduction success stories anyway. But I guess it's then you open yourself up to this problem of like people may stumble onto things that look a lot like inoculation prompting. And, you know, obviously we've got whatever 8 billion monkeys in the world that can prompt these things that the sort of infinite monkeys on infinite typewriters theory is like pretty closely approximated by how humanity at large is going to prompt a is. So any thoughts on inoculation prompting? And then I guess more broadly, I think you've had this feedback from time to time where people are like, well, when you prompt it like that, you know, you're going to get this. And I always feel like that is frustrating or misses the point because it's just my working model is like any prompt that could be written will be written. And you can't really excuse AI or, or certainly like, you know, you can't say we don't have a problem here just based on the fact that there was a prompt that you thought was, you know, maybe more suggestive than some other, you know, hypothetical prompt might have been. So what's your take on this sort of inoculation, prompting and and prompting discourse generally?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.