Evidence receipt / belief
Published · transcript-backedAxel Højmark: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation.”
Source trail
Everything needed to verify it.
- Speaker
- Axel Højmark
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…I suppose there's a question of do we need to have some kind of reward seeking for intelligence? Because in something like AlphaGo Zero, the domain is sufficiently rigid that it can be basically coded explicitly as part of the algorithm. But for a true intelligent process, it needs to have some kind of abstract notion of reward seeking, because usually the developers of the system won't know enough about the problem to specify how to solve it. So, is it I don't know if your prescription is to get rid of reward seeking entirely or just to make it more aligned. Yeah. Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation. From a from a cognitive perspective, whether things have goals or not, understanding things as if they have goals is very, very important. It's the way that we understand each other. Like, you know, and it's the best form of compression in language for me to identify a person or an organisation, and I say, this is their agenda because I can then communicate that to you. And increasingly, we're using these these goals based language to understand and talk about language models. But it's more than that. When we look at their chain of thought, it looks like they are using intent based abstractions to navigate the world. And and and isn't that interesting? Right? Because, you know, initially, language models, they were just dealing with instance by instance examples, and then we train them on code. And now they can look at a Python function description, and they can generalize many different inputs and outputs. And now we're seeing this tendency for them to use intent based language to understand the world. And this is great from a cognitive sophistication point of view, but the problem is with increasing levels of abstraction, you have increasing problems in interpretation.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.