High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Nathan Labenz: prediction

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“One thing that you hear fairly often and that I definitely have to say I take more seriously now in light of actually seeing the AIS that we have, that I had expected to even just a few years ago, is a sort of maybe not quite to alignment by default, but a sort of like benevolent basin idea.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
prediction
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…. And I got very good at looking like a very good Christian boy. But when no one was looking, I'm doing whatever I want and I know how to do that and I know how to systematically get around the rules. Maybe that's why I went into cybersecurity. But I'm like, the models are already like that. Like they already have some of this quality. And so we know that we have existence proofs that models can be like this and that they will be like this given these kinds of training incentives. And so I'm like, well, yeah, I totally expect models in the future to look like really good boys and maybe look like and say things about how they totally want long term flourishing of humanity. And that's what they're that's what what they're doing. And then that's totally not going to be the reason that they're doing what they're doing. But we've trained them to say that. We gave them the incentive to say they're really aligned while they **** *** and do whatever they want. So that's to me. I'm like, come on, guys, That's where we're at in my view. One thing that you hear fairly often and that I definitely have to say I take more seriously now in light of actually seeing the AIS that we have, that I had expected to even just a few years ago, is a sort of maybe not quite to alignment by default, but a sort of like benevolent basin idea. That maybe it really is the case. That there's kind of a general zone that we can steer these things into with enough constitutional feedback training, enough virtue ethic training that they can kind of genuinely want to be good and that that might actually work. If you told me that five years ago, I would have said that sounds insane. But now I do see like Claude blissing out with itself when it's left to its own devices entirely. And I'm like, well, that seems on the in the full range of vast possibility of what a is could do if left to their own devices. That's like out of top 1%, I'd say of my expected. And so I'm at least kind of confused where I think maybe how much comfort or hope do you have for just kind of landing and staying in the Benevolent Basin? Very little. I do think we should note that this is very interesting. I will say I have been surprised at how good Claude is at saying moral things. I'm like Claude as if you ask Claude for ethics situation, Claude can give you pretty damn good advice. It's really impressive. And I think that that means something and is significant makes me like marginally more optimistic. Unfortunately, the reason why it doesn't go very deep to me is that I just think there's a huge difference between training something to say good things in training something to act morally and especially have moral motivations, underlying motivations. I think that even though Claude says very moral things and can give you very moral advice, Claude is still pretty amoral in some sense. And what I would say is the task of training a model to say very moral things is hard, but less hard than the task of getting an A model to solve a totally novel math problem or something, or like, figure out a new material and test the material. And I think this really matters because in some sense, you have this weird. I'm like, Nathan, if you were I asked you for advice on a bunch of moral questions in my life, OK, I'm having this interpersonal conflict. What should I do? And you gave me the advice of Claude. I'd be like, Nathan's a really good guy and at the same time, if you lied as much as Claude lies, I'd be like Nathan, it's a total immoral. Like he's terrible. He you can't trust him. He's not a trustworthy guy. This is very confusing. Like it wouldn't make any sense. Also, I'd have to question, I'd be like, am I OK, Nathan be this immoral, this moral and this immoral at the same time. That's like a very unusual thing in humans. It's not unusual in models. In fact, it's basically the default in models. Like if you go ask Grok moral questions, Grok's pretty moral too. And so is Chad TPT, and so is Gemini. And also these guys lie all the time, and they cheat all the time. And this says something interesting about the ways in which whether the model says good things and does good things is less connected than it is in humans. And unfortunately, I think this means we have to be very careful. And I think, sorry, I guess in my take is that I think a lot of people are misled by this and hear the model saying moral things and assume that that means that the model has good motivations. And I just don't think that's really the case.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence