Evidence receipt / belief
Published · transcript-backedAxel Højmark: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL.”
Source trail
Everything needed to verify it.
- Speaker
- Axel Højmark
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…It's interesting as well that you're using the language of capabilities. I mean, because I I said intelligence before. So for for for me, my understanding of RL training is that it makes the models more intelligent. And for me, intelligence is the capability to acquire capability. So you put them in a novel situation, you know, the benchmarks went right up after o1. And and I felt it was because they now have this adaptation. So they've been trained to explore and to use what they know in a in a new way through some kind of composition and trial and error and and thinking and and whatnot. And and now it can adapt to novelty. So I I would say it's more intelligent, but but would you would you use the word intelligence or not? I definitely wouldn't mix in reward seeking and intelligence. I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL. The model is getting generally more capable at calling tools and doing tasks and so on, but it's also developing this specific type of reasoning and behavioural pattern that we call reward seeking, and both are sort of co occurring. I suppose there's a question of do we need to have some kind of reward seeking for intelligence? Because in something like AlphaGo Zero, the domain is sufficiently rigid that it can be basically coded explicitly as part of the algorithm. But for a true intelligent process, it needs to have some kind of abstract notion of reward seeking, because usually the developers of the system won't know enough about the problem to specify how to solve it. So, is it I don't know if your prescription is to get rid of reward seeking entirely or just to make…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.