High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Nathan Labenz: uncertainty

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“There's like a very direct incentive to dial that up and I don't know how we're going to figure out how to balance that.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
uncertainty
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, it's certainly not something we should be taking for granted. I can say that with 100% confidence. I've talked about this many times, but a formative experience for me was doing the GPT 4 Red team and using the helpful only model and just realizing how vast the space of AI mines really is and how easy it is for them to end up in a state that really violates our intuition for how people are going to be. As you said, much more correlated along different dimensions than minds in general have to be or that a is have to be. And so that's definitely something we should be keeping in mind a lot. I think another big interesting thing that's coming coming up on the alignment frontier is multi agent competitive world where mostly so far we've trained things to handle one thread, pursue 1 task for one user without too much in the way of dynamics. I'm sure you've followed and in labs work with their vending bench and things like that. It's been interesting to see recently that the most recent claws they've described as being ruthless. I guess it's kind of there's a comment here and then there's I think a question as well. The observation is that's kind of a Yikes. And I don't know, obviously I don't know all that's going on in quad training. But it sure seems like we are entering A regime where one of the very natural next things to do is going to be to train agents in competitive environments where they're supposed to make money or they're supposed to negotiate or they're supposed to represent interests in a world where other agents or entities have other interests. That seems like it's going to be a big Yikes because that the world itself just rewards deception. We see deception in nature all over the place. And there's a very fundamental reason for that, which is you can win by deceiving the other agents in your environment. So I'm really on the lookout right now for how are companies going to handle that if they want their AIS to be able to go out? And by the way, the economy is an adversarial environment, right? Like if you are naive and you go out into the world of suppliers and negotiations or whatever, and you take everything at face value and you don't try to push back a little bit. Or if, if you don't have some separation between your initial offer and your bottom line reservation price, then you're going to be the sucker that's taken advantage of. And nobody's going to want to use that AI to go out and do these sorts of things, right? We, I think there's this big dream, which I'm excited about of like a is taking search costs super low, kind of facilitating all these transactions that previously couldn't have happened because the transaction costs were too high to facilitate. But to do that well, they are going to have to have certain amount of at least minimal deception. And it seems like we're we're already kind of seeing it arise through sort of whatever prior and accident and there's now seemingly very already too soon. certain amount of at least minimal deception. And it seems like we're we're already kind of seeing it arise through sort of whatever prior and accident and there's now seemingly very already too soon. There's like a very direct incentive to dial that up and I don't know how we're going to figure out how to balance that. So a comment that is very. Important here, which is in nature, deception is highly incentivized in many, many cases. And it's very interesting because you get deception in a system that doesn't have a mind. Like I've been learning about flowers recently. I have an evolutionary biology background, but I was studying bats and monkeys my undergrad and not. So it's all animals. I didn't really study plants much at all. So recently I've been getting into plants. Plants are fascinating. A notable feature of plants. No mines, like they have some sensory capacity, but it's really the evolutionary process where you see deception show up in plants. And so you have all these different orchids, there's thousands and thousands of orchid species. And many of them are extremely deceptive. They will basically create this like shape that looks exactly like a bee or a wasp, some type of insect. And that insect will go and try to mate with the orchid. And this is to pollinate the orchid. But like the bee doesn't get anything out of it. It like it only. In fact, it's like it's parasitic. It's like the bee is foregoing reproductive activity, reproductive opportunities, It's hoping to *** **** and it's not. It's getting a flower instead. And and then it goes and does that with another orchid flower of the same species. And then the orchid gets pollinated and the bee has to go find an actual mate. So, and there's many, many instances of this with many different insects across many different flowers. And you're like, no, that's just natural selection. Just founded a good deceptive strategy that worked here. And I think what this implies, which is like what you said, is that unfortunately deceptive deception is a very natural strategy. And I, and I think people get this wrong. I think a lot of people are like, oh, humans are uniquely sinful and fallen. And so the AIS won't be deceptive unless we like, unless they learn from us or like we teach them that. I'm like, no, that's not how it works. Like, like, unfortunately deception is very common in nature and it's a very natural strategy. And one of the things that makes humans unique is that we have managed to create a value of honesty and we have managed to create culture and coordination around let's not do the natural deceptive thing. Let's like try to rise above and have better coordination. And like, I think that I just think that when I think we have a lot of evidence for this, the natural basin that models will fall into is one that's extremely deceptive. And we need to figure out a way to get the, the models into a, the basin of, of honesty and coordination that, that humans have sometimes found. And that's going to be a challenge. And I think it's possible. I think I really do believe in a future where we could have AIS that are mediating human interaction in a way where we don't have wars anymore, right?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence