Evidence receipt / belief
Published · transcript-backedAxel Højmark: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for.”
Source trail
Everything needed to verify it.
- Speaker
- Axel Højmark
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Now Fable has come out, and Fable presumably has just been RL tuned to oblivion, And it is so overeager. It's so adaptive because intelligence is about adaptivity. It's about, like, not what the model knew before. It's about, I'm in a novel situation. I can use all of these tools, and I can just do all of these crazy things. So the other day on our Discord server, Fable deleted a 100 messages from Wendy. I'm very sorry about that, Wendy. I was doing some financial analysis, and, you know, I put I put a thing up in in Notion. And then it said under the image, also uploaded to archive.mlst.ai. And I was like, what the fuck? You Do know what I mean? And so this Fable is just so overeager. It's so adaptable, and it's weird because this we thought this was what we wanted. Right? And this is actually gonna be a nightmare, isn't it? Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for. So 1 of the big safety issues is that, yeah, when you're in deployment, it might be unclear what's graded. It's plausible that reward seeking might generalize worse. So the Fable system card came out, and 1 thing we found in there is using natural language autoencoders, they looked at this reasoning about the grading process, awareness of grading, and what they actually found in that paper is that across more training, this also goes up. And this basically corroborates the same finding that we have but on a completely different model family with a completely different method and that to us was pretty significant because it suggests that this is a broader trend that doesn't just affect, like, single model, but, like, probably most models at certain amounts of, like, training compute.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.