Evidence receipt / belief
Published · transcript-backedAxel Højmark: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“I think well, 1 thing is just dedicating more resources into, like, investigating this phenomenon, getting better ways of measuring.”
Source trail
Everything needed to verify it.
- Speaker
- Axel Højmark
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yes. And it is conceivable that this process is accelerating over time, especially, like, everyone is now talking about continual learning. You know, because right now, at least we have a centralized model. Right? So so we have we have red teaming. We have, you know, frontier companies building these these models, testing them. It's gonna become far more diffused and decentralized and and all of this kind of stuff. I mean, what what is your prescription and what is your prognosis? Yeah. I think well, 1 thing is just dedicating more resources into, like, investigating this phenomenon, getting better ways of measuring. Like, the 1st thing when you want to make an intervention on something is having a really good measurement of it, and that's essentially the vision we had for the project of trying to establish that, but we need far more work making it more robust and applying it to even more frontier models and so on. And then 2nd, like, just making it a default thing that the labs track. They all, like, track how reward seeking are the different various measures. Anytime the model in training faces a trade off between doing what it believes is intended versus doing what it believes is rewarded, it gets by definition rewarded for the cases where it ignores the actual intent. So the more RL we throw at models, the more we should expect this tendency to go up. The optimization pressure away from thinking in human ontologies is of course there. Right? If we want to create superhuman intelligence, then superhuman capabilities even in narrow domains, it probably requires that the model also uses internal representations that are not natural to humans. We would expect that this drifts away from the pretraining representations. We do already see this. Right? What are the odds that chain of thought in legible, normal English is just the perfect language for reasoning? So even if there weren't strong length penalties or anything like that, you would expect some ontological drift over time in so so tokens start meaning something else or they start acquiring double meanings and so on. And then the 2nd thing that you said is that basically there could be a new optimization loop that operates at a different different speed. And I think in principle it could happen that once the majority of the context that agents read and encounter is created by other agents, you have something akin to memetic evolution inside there. Right? Like, just as a super dumb example, if there was a sentence that says copy me a 100 times and agents for whatever reasons there are prompt injectable by this thing, you you would propagate it throughout the agent population. We're now in the regime where slowly more and more of the state that agents encounter is agent generated, so there there would be a 2nd optimization loop of the yeah. It would be analogous to, like, the culture. Although, to be fair, this is not the type of optimization pressure that we think about a lot or that we that we directly work on.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.