High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Alexander Meinke: belief

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“Although, I think, like, in principle, the the math example and drug discovery, I can totally see models being very useful in those domains without goal language being applicable.”

— Alexander Meinke

Source trail

Everything needed to verify it.

Speaker
Alexander Meinke
Attribution
Verified speaker
Claim type
belief
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…This sort of disagreement was a lot more applicable and yeah. Applicable back in the early days when you had these very, very small simple transformers. They were, like, statistical pattern matchers of predicting the next token and so on, and they're sort of pushing the frontiers on, like, math and also drug discovery and things like this. Yeah. Yeah. I think I think a model that says it's, like, just like a a pattern matcher just doesn't fit the facts anymore. When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example. Although, I think, like, in principle, the the math example and drug discovery, I can totally see models being very useful in those domains without goal language being applicable. So for example, like, AlphaFold 3 or something, there I would not think that the goal of solving protein folding would be appropriate in that case. Goal directed language or, like, intent or whatever becomes useful when the model has ontologies that sort of, like, broadly represent the the world. They have a world model in some sense. And secondly, they are goal directed in in the sense that you can best describe their the the outcomes of their behavior by saying this is the thing that they're trying to achieve. And it helps you sort of predict the behavior a little out of distribution and so on. And yet, I think this doesn't apply to, yeah, chess engines or something. would would you guys field the question about what the purpose of your organization is?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence