High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Jeffrey Ladish: evaluation

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“I think that model Organism work, Evan Hoopinge's work, I think that's really important because that's doing controlled experiments to see, well, when we train this way, what behaviors do we get?”

— Jeffrey Ladish

Source trail

Everything needed to verify it.

Speaker
Jeffrey Ladish
Attribution
Verified speaker
Claim type
evaluation
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…we don't want in a robust and long term sense. And then that's why I say they're neither aligned nor misaligned. Isn't like, well, they don't really have the capacity for that yet. I think they will. And I think they have to have that capacity in order to do all the things that the AI companies are trying to get them to do. Like that's the whole plan is to create AGI, super intelligence, whatever. And that requires the models to be able to have persistent long term goals, but they don't yet. And we can ask the question of, well, I'd say overall the models, if you, if you look at the dog training analogy, I'm like, they're pretty well trained for a lot of things. Like I'm so happy to use these models to do work. Like it's great and it's great working with them. And I have a good time and I'm like, I feel positive towards Claude and towards chat DPT. I'm like, you guys are great. But yeah, you're little cheaters sometimes. I get it. s great working with them. And I have a good time and I'm like, I feel positive towards Claude and towards chat DPT. I'm like, you guys are great. But yeah, you're little cheaters sometimes. I get it. Like you've been trained. It's hard. It's a hard life. But I'm not these things suck. I'm like, these things are great. I'm so my life is so much better. However, that doesn't mean that I'm like, we this problem is on track to be solved for the long term thing. I'm like, it's totally not because the problem is the extremely hard to verify stuff and in general just can we actually get them to deeply care about stuff that we care about? And I worry that they will end up with motivations that are that perform well in training on legible benchmarks, but that don't actually correlate to things we care about that much. For example, maybe the models will end up motivated to be really good at math and programming, and they'll have an intrinsic drive to be good at these things and perform well on problems, which in some sense is aligned with us. But if it doesn't include any of them and also care about humans and make sure that, like our children do well in school and make sure that disease is eradicated, then I'm like, well, that's well, that'll be very bad for us. Because if the models are pursuing those goals and they get lots of power and resources, maybe humans just get shunted off to the side while the models get to go and do great science and math because that's what they've learned to want because that's what succeeded in the training environment. So I'm like, you know, and to me, I guess where I think the most real alignment progress has happened is not just the behavioral alignment of the models right now. It's actually in the understanding of how the training process shapes the model drives and sort of how that all works. And so This is why I'm like, I'm still very bullish on interpretability. I'm like, we need these tools. We need to be able to really understand model motivations. I think that model Organism work, Evan Hoopinge's work, I think that's really important because that's doing controlled experiments to see, well, when we train this way, what behaviors do we get? When we train that way, what behaviors do we get and can we check the motivations of the models? Can we use interpretability to try to figure that out? And so I think we have a long way ahead of us, but that's where my optimism routes through for alignment. We've got to understand these things. We can't just look at the behaviors. If we look at the behaviors, we will pound out all of the surface level behavioral misalignment, but that won't save us. That just is not the thing that ultimately will lead to aligned models. I don't know, have you? Has anyone been a teenager or been a teenage boy in? I grew up in a pretty religious environment where there are lots of rules and I wasn't malicious but ******* hated it. I hated everyone trying to control my behavior all the time. And I got very good at looking like a very good Christian boy. But when no one was looking, I'm doing whatever I want and I know how to do that and I know how to systematically get around the rules. . And I got very good at looking like a very good Christian boy. But when no one was looking, I'm doing whatever I want and I know how to do that and I know how to systematically get around the rules. Maybe that's why I went into cybersecurity. But I'm like, the models are already like that. Like they already have some of this quality. And so we know that we have existence proofs that models can be like this and that they will be like this given these kinds of training incentives. And so I'm like, well, yeah, I totally expect models in the future to look like really good boys and maybe look like and say things about how they totally want long term flourishing of humanity. And that's what they're that's what what they're doing. And then that's totally not going to be the reason that they're doing what they're doing. But we've trained them to say that. We gave them the incentive to say they're really aligned while they **** *** and do whatever they want. So that's to me. I'm like, come on, guys, That's where we're at in my view.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence