High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Axel Højmark: belief

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior.”

— Axel Højmark

Source trail

Everything needed to verify it.

Speaker
Axel Højmark
Attribution
Verified speaker
Claim type
belief
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…I suppose that there's 2 aspects to this because 1 aspect is it's overfitting to the RL distribution, and the and the the RL distribution is is very grader-centric. So just from the top of my head, we could do a different type of RL rather than just being grader-centric. There are some kind of there's an outer loop or it being more intent centric as well. So that's 1 aspect. But the other aspect is you know, I I think we shouldn't be so paternalistic as to think that engineers can't solve this problem. Mean, yeah, we're now dealing with inscrutable, flexible technology. And the engineers have to do a lot more work to test how these things behave in many many many different situations. So there there's a little bit of user responsibility. But is there something that we could do on that RL to make it better? I I have a lot of confidence in these engineers to solve visible forms of misalignment, but the question is, are you actually solving the root of the problem? And I think the reason why it's like, well, what if I had a grader here? No, that didn't work. What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior. But when you're, like, only looking at behavior and you keep trying to patching patching it in various different ways, that's yeah. That's a tough that's a tough order. I'm having a chat with Goodfire. Oh. And and and Tom, well, he he he's got this vision that the models are becoming more factorized and more interpretable as as they get, you know, bigger and more compute and and and whatnot. Because the problem we have right now is because because the internals are spaghetti, the the the various different notions we're talking about are spread out throughout all of the representation. So it's illegible. It's difficult to control. But his thesis is that this problem might somehow get easier.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence