Evidence receipt / belief
Published · transcript-backedAxel Højmark: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example.”
Source trail
Everything needed to verify it.
- Speaker
- Axel Højmark
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah, absolutely. I mean, example, if I speak about Greta Thunberg, the 1st thing that comes to your mind is her environmentalist agenda. It's an incredible form of compression. And it doesn't surprise me in the slightest that language models have converged on that as an ideal compressed form of representation. You know? Maybe maybe it's it's a natural convergent thing. But the thing is, I'm I'm in I'm in 2 minds. Because when I saw Ryan Greenblatt talking about scheming and, you know, like, we've been using words like lying and deception. And it feels like we are implying that these things have a mind. This sort of disagreement was a lot more applicable and yeah. Applicable back in the early days when you had these very, very small simple transformers. They were, like, statistical pattern matchers of predicting the next token and so on, and they're sort of pushing the frontiers on, like, math and also drug discovery and things like this. Yeah. Yeah. I think I think a model that says it's, like, just like a a pattern matcher just doesn't fit the facts anymore. When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example. Although, I think, like, in principle, the the math example and drug discovery, I can totally see models being very useful in those domains without goal language being applicable. So for example, like, AlphaFold 3 or something, there I would not think that the goal of solving protein folding would be appropriate in that case. Goal directed language or, like, intent or whatever becomes useful when the model has ontologies that sort of, like, broadly represent the the world. They have a world model in some sense. And secondly, they are goal directed in in the sense that you can best describe their the the outcomes of their behavior by saying this is the thing that they're trying to achieve. And it helps you sort of predict the behavior a little out of distribution and so on. And yet, I think this doesn't apply to, yeah, chess engines or something. would would you guys field the…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.