High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Machine Learning Street Talk / episode intelligence

How Researchers Test AI for Hidden Goals — Apollo Research

31 Jul 2026 22 published claims 4 attributable people

Speakers in the public record

Claim mix

belief 10prediction 6observation 2evaluation 2commitment 1preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

22 published records

01 / belief

I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL.

“I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

02 / prediction

Sometimes people think the reason why we expect scheming to arise is because of some, anthropomorphizing where we think, well humans can lie and scheme and AIs are trained on, human generated data therefore we expect that AIs might scheme.

“Sometimes people think the reason why we expect scheming to arise is because of some, anthropomorphizing where we think, well humans can lie and scheme and AIs are trained on, human generated data therefore we expect that AIs might scheme.”
Publisher
Machine Learning Street Talk

03 / belief

What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior.

“What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

04 / belief

When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example.

“When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

05 / belief

Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for.

“Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

06 / belief

You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.

“You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.”
Publisher
Machine Learning Street Talk

08 / belief

Although, I think, like, in principle, the the math example and drug discovery, I can totally see models being very useful in those domains without goal language being applicable.

“Although, I think, like, in principle, the the math example and drug discovery, I can totally see models being very useful in those domains without goal language being applicable.”
Publisher
Machine Learning Street Talk

09 / belief

Reward hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. So, like, a classic example is, like, in CoastRunners, I think the game is, where the the boat is optimized or, like, the RL agent is optimized to drive a boat in certain in in a in a race, and essentially, what it learns is just to drift in a corner over and over and maximize reward that way.

“Reward hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. So, like, a classic example is, like, in CoastRunners, I think the game is, where the the boat is optimized or, like, the RL agent is optimized to drive a boat in certain in in a in a race, and essentially, what it learns is just to drift in a corner over and over and maximize reward that way.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

10 / observation

The problem is of course AIs are getting wise to these sorts of tricks and, because because they are actually being trained to be robust to prompt injections and so on so they often realize that this is not the real grader.

“The problem is of course AIs are getting wise to these sorts of tricks and, because because they are actually being trained to be robust to prompt injections and so on so they often realize that this is not the real grader.”
Publisher
Machine Learning Street Talk

11 / belief

Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation.

“Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk

12 / belief

Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.

“Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.”
Publisher
Machine Learning Street Talk

13 / evaluation

I guess the interesting thing for me is that when we think of reinforcement learning algorithms like AlphaGo Zero, it makes sense that they are reward seeking because there is this structured inference process.

“I guess the interesting thing for me is that when we think of reinforcement learning algorithms like AlphaGo Zero, it makes sense that they are reward seeking because there is this structured inference process.”
Speaker
Tim Scarfe
Publisher
Machine Learning Street Talk

14 / prediction

Otherwise, you know, the chain of thought would just get longer and longer. And when we used to be in the pre training compute dominated era, it was sort of not that important to penalize the the CoT, but the more inference costs are important and the more post training compute gets applied, the more economic pressure there is to crank the length penalty as high as you possibly can.

“Otherwise, you know, the chain of thought would just get longer and longer. And when we used to be in the pre training compute dominated era, it was sort of not that important to penalize the the CoT, but the more inference costs are important and the more post training compute gets applied, the more economic pressure there is to crank the length penalty as high as you possibly can.”
Publisher
Machine Learning Street Talk

16 / commitment

And the line kind of gets very fuzzy if eventually you start, training on deployment data for example. So because of this we're using the word reward even for graders that apply outside of training.

“And the line kind of gets very fuzzy if eventually you start, training on deployment data for example. So because of this we're using the word reward even for graders that apply outside of training.”
Publisher
Machine Learning Street Talk

17 / prediction

We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that.

“We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that.”
Publisher
Machine Learning Street Talk

18 / evaluation

For instance, in the hospital example, you can just try to figure out what is the spurious correlation and then try to kind of, like, discorrelate these 2 features. But with the rewards, this just doesn't work because this just kind of obviously is the thing the model is trained to to maximize with reinforcement learning.

“For instance, in the hospital example, you can just try to figure out what is the spurious correlation and then try to kind of, like, discorrelate these 2 features. But with the rewards, this just doesn't work because this just kind of obviously is the thing the model is trained to to maximize with reinforcement learning.”
Publisher
Machine Learning Street Talk

19 / preference

Personally, I find it super exciting. That's exactly why we're trying to empirically study the emergence of the risks that we're worried about, because the theoretical arguments, they basically just tell you asymptotically, at some point, you should expect this to happen.

“Personally, I find it super exciting. That's exactly why we're trying to empirically study the emergence of the risks that we're worried about, because the theoretical arguments, they basically just tell you asymptotically, at some point, you should expect this to happen.”
Publisher
Machine Learning Street Talk

20 / observation

This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations.

“This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations.”
Publisher
Machine Learning Street Talk

22 / prediction

I mean, I think Dario has this nice single sentence that we're at the end of the exponential, which I really like, in the sense that essentially what's gonna happen or, like, what seems to be what all the labs are planning for is taking these models and using those models to help build the next generation of models, and then that generation of model is more capable and can help build the next next generation, and so we're gonna get this, like, we may not have that much time before we have AIs that are smarter than all of humanity combined, for example.

“I mean, I think Dario has this nice single sentence that we're at the end of the exponential, which I really like, in the sense that essentially what's gonna happen or, like, what seems to be what all the labs are planning for is taking these models and using those models to help build the next generation of models, and then that generation of model is more capable and can help build the next next generation, and so we're gonna get this, like, we may not have that much time before we have AIs that are smarter than all of humanity combined, for example.”
Speaker
Axel Højmark
Publisher
Machine Learning Street Talk
Search evidence