High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Jeffrey Ladish: uncertainty

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“Maybe not 100%, but we could even see where in training this type of behavior comes from. And I'm like, well, hell yeah, I want to celebrate that success because I'm, I don't know.”

— Jeffrey Ladish

Source trail

Everything needed to verify it.

Speaker
Jeffrey Ladish
Attribution
Verified speaker
Claim type
uncertainty
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Well, and you're seeing things like in the the meter report that just came out on, on evaluating risks of losing control across the board. All of these models doing these evaluations, the majority of the time and effort they spent was on figuring out how to get the models not to cheat or how to evaluate their performance in difficult tasks. When the models have a strong inclination, the more difficult to task, the more likely they are to cheat. And that's very telling to me. And also the models are often in their chain of thought saying like, I'm going to cheat here. Oh, I can totally like hack this and the models know what they're doing, but it's just very hard to incentivize the models to not cheat in cases where the task is hard to verify. And so I think that that's there's this whole question about how is alignment going? How is the science of alignment going? And I think the good news is, is that it seems like the models are not scheming in a long term sense. It seems like the models have not yet developed a survival drive. It seems like they're not pursuing misaligned objectives in a strategic long term sense. And that's great news. The bad news is it seems like models are persistently misaligned on the stuff that they're actually good at, especially as the stuff is harder and harder to verify. And the reason I think this is important, So really difficult coding challenges, really long time horizon tasks. And in some sense, if you could say, well, the labs sure do have an incentive to get them to be better at those long time horizon tasks that they're currently cheating at. So they're naturally they're going to have to work on alignment here. And that's true. But it's really a problem if the things we need the models to be aligned on are the hardest to verify. So for example, if you need the models to be aligned on the long term trajectory of humanity, you know, if it's the thing on the 20 year time scale, 50 year time scale, if you need them to be really aligned on that, say if they have a lot of power and control, then that's going to be extremely hard to verify. And that's going to be the thing that they're most likely to be misaligned on, which is the thing we most care about. So that's where I'm like, I don't feel good about the current alignment progress. I feel good about We're learning a lot and that's great. The interpretability work coming out of Anthropic I think is excellent. The thing we're just talking about was blackmail. I remember going to an AI conference with a lot of Anthropic and open AI researchers and I'm like, can we talk about blackmail? I don't actually understand exactly why this is happening. And it's like, it's very high profile. Like I've talked about it in a documentary. It's been talked about by people high up in the admin. We should know exactly why this happens, right? And then Anthropic did a bunch of, I think pretty good interpretability work. t in a documentary. It's been talked about by people high up in the admin. We should know exactly why this happens, right? And then Anthropic did a bunch of, I think pretty good interpretability work. And they're like, hey, we have a much better idea of why this happens. Maybe not 100%, but we could even see where in training this type of behavior comes from. And I'm like, well, hell yeah, I want to celebrate that success because I'm, I don't know. I was like calling for it and being like, we need this. And then researchers came through and they're like, hey, we actually now can trace where this behavior comes from. It's this persona misalignment thing here how it works. And I'm like that. I mean, that is exactly what we need. If we can deeply understand how the training process shapes the drives and motivations of the models, then we might have a shot at actually crafting those drives and motivations intentionally so that we can get this longer term alignment. And we'll know if that's working at all if. The models start to suddenly become gradually become very aligned on these hard to verify tasks. If they stop cheating at hard to verify tasks, that doesn't mean that our job is done, but it means that we're making significant progress at some piece of the hard part of the problem. Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence