High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Jeffrey Ladish: prediction

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“Let's like try to rise above and have better coordination. And like, I think that I just think that when I think we have a lot of evidence for this, the natural basin that models will fall into is one that's extremely deceptive.”

— Jeffrey Ladish

Source trail

Everything needed to verify it.

Speaker
Jeffrey Ladish
Attribution
Verified speaker
Claim type
prediction
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…certain amount of at least minimal deception. And it seems like we're we're already kind of seeing it arise through sort of whatever prior and accident and there's now seemingly very already too soon. There's like a very direct incentive to dial that up and I don't know how we're going to figure out how to balance that. So a comment that is very. Important here, which is in nature, deception is highly incentivized in many, many cases. And it's very interesting because you get deception in a system that doesn't have a mind. Like I've been learning about flowers recently. I have an evolutionary biology background, but I was studying bats and monkeys my undergrad and not. So it's all animals. I didn't really study plants much at all. So recently I've been getting into plants. Plants are fascinating. A notable feature of plants. No mines, like they have some sensory capacity, but it's really the evolutionary process where you see deception show up in plants. And so you have all these different orchids, there's thousands and thousands of orchid species. And many of them are extremely deceptive. They will basically create this like shape that looks exactly like a bee or a wasp, some type of insect. And that insect will go and try to mate with the orchid. And this is to pollinate the orchid. But like the bee doesn't get anything out of it. It like it only. In fact, it's like it's parasitic. It's like the bee is foregoing reproductive activity, reproductive opportunities, It's hoping to *** **** and it's not. It's getting a flower instead. And and then it goes and does that with another orchid flower of the same species. And then the orchid gets pollinated and the bee has to go find an actual mate. So, and there's many, many instances of this with many different insects across many different flowers. And you're like, no, that's just natural selection. Just founded a good deceptive strategy that worked here. And I think what this implies, which is like what you said, is that unfortunately deceptive deception is a very natural strategy. And I, and I think people get this wrong. I think a lot of people are like, oh, humans are uniquely sinful and fallen. And so the AIS won't be deceptive unless we like, unless they learn from us or like we teach them that. I'm like, no, that's not how it works. Like, like, unfortunately deception is very common in nature and it's a very natural strategy. And one of the things that makes humans unique is that we have managed to create a value of honesty and we have managed to create culture and coordination around let's not do the natural deceptive thing. Let's like try to rise above and have better coordination. And like, I think that I just think that when I think we have a lot of evidence for this, the natural basin that models will fall into is one that's extremely deceptive. And we need to figure out a way to get the, the models into a, the basin of, of honesty and coordination that, that humans have sometimes found. And that's going to be a challenge. And I think it's possible. I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right? ing to be a challenge. And I think it's possible. I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right? And like, because we can, we have found better ways to resolve conflicts because we have these like smarter, more powerful arbiters who are able to, who are not like authoritarian controlling us, but who are able to like, help mediate conflicts in ways that are actually positive sum for people. But I think we really have to get through this basin of extremely deceptive behaviour in order to get there. And like as you said, once you're like, we are already encountering models that are cheating like a lot in cases where and it's just on a computer. It's just like in programming tasks. Once we get into like economic tasks where like you said, there's even much more incentive for deception, then I'm like that's playing on hard mode. And if you go even further than that, if you try to make war clod, we're trying to make Claude that will go infiltrate the CCP and like live out in like the Chinese tech company servers and like spy on them and sabotage them on on its own without oversight or supervision. I'm like, Oh my God, that is extreme hard mode. How do you align that system so it like will **** with your adversaries, but be nice to you. Like that's a very tricky and we know from you like there's plenty of, you know, double agents or double agents turned triple agents in, in human spycraft history working with human minds. We do understand somewhat well. So yeah, I think it's that's I think that's we have some some real challenges ahead as we move into more competitive domains. And this is a thing that a policy we think about a lot, which is like it's not just that we have to solve alignment, we have to solve alignment given these competitive pressures. And I don't know.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence