Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
“I think a general sketch would be like 03 might be the most misaligned model that was ever released to the public. It seemed like it was right in that tween zone where RL had really scaled up and some of these problems were starting to show up, and since then there's been a bunch of work to try to reduce them.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 24 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr Before we go into your next result, how would you characterize progress as it has unfolded recently? I think a general sketch would be like 03 might be the most misaligned model that was ever released to the public. It seemed like it was right in that tween zone where RL had really scaled up and some of these problems were starting to show up, and since then there's been a bunch of work to try to reduce them. My mental model usually is like with each scale up we seem to get new kinds of bad behaviors. Some have argued to me that really it's just all reward hacking and different flavours. And I take that point. But then I still see qualitative differences in the shapes of the behaviour. So I'm not quite sure how I should be thinking about that. But it seems like we do see a new patterns of behaviour emerge with each capability advance, which obviously corresponds to another scale up and sometimes new techniques and training. And then as those arise, they come to prominence and then they get pushed down. And I think they're mostly getting pushed down by like a combination of just training against them. Meaning, OK, here's some examples of where we've had problems. Let's show the model what's good to do and maybe do some, maybe do a little supervise, maybe do a little reinforcement to try to tamp those behaviors down. Maybe we also do a little investigation into was there something in training data that we can specifically maybe not fully chain this back to, but at least to some degree we can do that. And then we can alter the mix or we can filter some stuff out. And it usually seems like it goes down by like in the next generation, 2 thirds, 80% never goes to 0. But then we seem to move on. And so it seems like we have a lot of these things that are still at a low level and that's for now a tolerable situation. I mean, I'm certainly still using models all the time. Is there any would you the same story you see playing out or do you would you say anything differently about what what is actually happening and how successful it is? Well, I think it's pretty key to look at what is, what are the alignment problems we're trying to solve. So I would say the current models are pretty amoral. Like they're not, I don't think they're aligned or misaligned. I think that they don't have the capacity to be either in some sense. And by that I mean, if we're talking about trying to align AGI or super intelligence, you know, even human level intelligence, where you're talking about an agent that can run a company or an agent that can run a political campaign. Once you have an agent that there's a real sense in which is it a line? Does it miss the line? Is it going to screw people over? Is it going to try to cause effects in the world that hurt other people or help other people? Like there's a real, there's real stakes there and the real steering process that's trying to go towards one of these things. And I'm like, I don't think models are not capable of that right now. Like they don't have the time horizon. If they only have a 12 or 24 hour time horizon, they can't really steer reality towards some particular outcome with humans overall are way better off, worse off. That being said, they certainly do 'cause effects in the world and they have some awareness of this. And there's a question of are they following instructions? Are they following instructions as the developers intended, etcetera. And this has real safety implications, right? Like, so I don't know, you might think of this as as analogous to is the dog well trained, right? And then there's a separate thing, which is this, the dog actually care about you or something. And I think in the case of dogs, because animals that have the scenario feedback loop and training, the dog also cares about you in the long term sense. So in that sense, dogs are aligned in a real way, in a way I think that models aren't. But if you look at just the immediate, will the dog bite you? Or will it suddenly freak out and do a bunch of things you don't want? Or will the dog wait until you're out of the room and then jump onto your table and eat all your food? That's mainly the alignment problems that we're working on right now. And the more aligned the model, the more useful it's going to be because it's just like, it's not, you're not very happy with your dog if it, if it perfectly follows all of your, you know, rules and then you leave the room and it jumps on the table and eats all the food. And so I think we're currently in the case where we're struggling with those types of behaviours where it's like if the task is hard to verify, the model will often cheat or often fake work or whatever. But I do want to distinguish between that type of misalignment and the type of misalignment where the models pursuing a goal that we don't want in a robust and long term sense. And then that's why I say they're neither aligned nor misaligned. Isn't like, well, they don't really have the capacity for that yet. I think they will.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.