← All source episodes Machine Learning Street Talk / episode intelligence
How Researchers Test AI for Hidden Goals — Apollo Research
31 Jul 2026 22 published claims 4 attributable people
Speakers in the public record
Claim mix
belief 10prediction 6observation 2evaluation 2commitment 1preference 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
22 published records
“I mean, it's like if you're very intelligent, you're probably you're like a better reward seeker, everything else, and you have like, a stronger predictive model of the environment and and so on, and I think both are increasing during RL.”
- Publisher
- Machine Learning Street Talk
“Sometimes people think the reason why we expect scheming to arise is because of some, anthropomorphizing where we think, well humans can lie and scheme and AIs are trained on, human generated data therefore we expect that AIs might scheme.”
- Publisher
- Machine Learning Street Talk
“What if I had a grader here? It's all, I think, it's because we're in, like, an outcome based paradigm, and I think our alignment science is just very nascent, and we don't have a great understanding of model internals, and if we had, like, a perfect telescope and we could look into the model and see, here's the reward seeking, then maybe we would have a way better chance at essentially eliminating this behavior.”
- Publisher
- Machine Learning Street Talk
“When you look at, for example, Claude, it will, in a bunch of different contexts, try to steer away from giving the user instructions for building a bomb, for example, and it will, like, refuse and maybe try to give other suggestions, and I think when you see this, like, concrete pattern happening in many instances, it is like a useful model, predictive model, to call that a goal of not giving the user bomb instructions, for example.”
- Publisher
- Machine Learning Street Talk
“Like, for the normal user, I would guess that they will keep patching the reward function, they will keep making it closer and closer to the like, avoiding the types of annoyances that you had, and then the problems will seem to go away, but either they will still probably be, like, unverbalized representations of this, and especially, I think what's important is how does reward seeking generalize to context where it's not clear what it's graded for.”
- Publisher
- Machine Learning Street Talk
“You know, like, we wouldn't be claiming that, like, the model becomes more reward seeking because of know what I mean? It's like, there's something specific about, like, how you train the model that I would say makes it more reward seeking.”
- Publisher
- Machine Learning Street Talk
“I think well, 1 thing is just dedicating more resources into, like, investigating this phenomenon, getting better ways of measuring.”
- Publisher
- Machine Learning Street Talk
“Although, I think, like, in principle, the the math example and drug discovery, I can totally see models being very useful in those domains without goal language being applicable.”
- Publisher
- Machine Learning Street Talk
“Reward hacking is when you find unintended tricks or hacks or unintended solutions to problems that essentially developers didn't intend. So, like, a classic example is, like, in CoastRunners, I think the game is, where the the boat is optimized or, like, the RL agent is optimized to drive a boat in certain in in a in a race, and essentially, what it learns is just to drift in a corner over and over and maximize reward that way.”
- Publisher
- Machine Learning Street Talk
“The problem is of course AIs are getting wise to these sorts of tricks and, because because they are actually being trained to be robust to prompt injections and so on so they often realize that this is not the real grader.”
- Publisher
- Machine Learning Street Talk
“Think you you said at 1 point, like, we need reward seeking for these type of for these types of behavior, and I think, actually, you could go entirely without that, so you could have a model that thinks about how I'm graded and then smiles for that, or you could have 1 that's what are my developers' intent, what is my user intent, oh, they want me to look for files and so on, And so you can see the same behavior, same looking files, solving the task, but it's for entirely different internal motivation.”
- Publisher
- Machine Learning Street Talk
“Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.”
- Publisher
- Machine Learning Street Talk
“I guess the interesting thing for me is that when we think of reinforcement learning algorithms like AlphaGo Zero, it makes sense that they are reward seeking because there is this structured inference process.”
- Publisher
- Machine Learning Street Talk
“Otherwise, you know, the chain of thought would just get longer and longer. And when we used to be in the pre training compute dominated era, it was sort of not that important to penalize the the CoT, but the more inference costs are important and the more post training compute gets applied, the more economic pressure there is to crank the length penalty as high as you possibly can.”
- Publisher
- Machine Learning Street Talk
“1 is that if models are misaligned, then this is worse when they are more capable.”
- Publisher
- Machine Learning Street Talk
“And the line kind of gets very fuzzy if eventually you start, training on deployment data for example. So because of this we're using the word reward even for graders that apply outside of training.”
- Publisher
- Machine Learning Street Talk
“We don't yet understand, like, how or why this exactly happens, but I think 1 mental model that I use a lot, which helps me to think about this is, say you're a language model, and you're being trained to complete some sort of task, like, I don't know, sorting the files on a laptop, and you're being rewarded for that.”
- Publisher
- Machine Learning Street Talk
“For instance, in the hospital example, you can just try to figure out what is the spurious correlation and then try to kind of, like, discorrelate these 2 features. But with the rewards, this just doesn't work because this just kind of obviously is the thing the model is trained to to maximize with reinforcement learning.”
- Publisher
- Machine Learning Street Talk
“Personally, I find it super exciting. That's exactly why we're trying to empirically study the emergence of the risks that we're worried about, because the theoretical arguments, they basically just tell you asymptotically, at some point, you should expect this to happen.”
- Publisher
- Machine Learning Street Talk
“This model is in fact changing its behavior more towards the grader at the cost of the other authority such as the user. And so what that means is we observe after SDF, so after synthetic document fine tuning, that the model will explicitly reason about what is being rewarded and then take that action in various situations.”
- Publisher
- Machine Learning Street Talk
“For the 1st time now everyone's thinking about sovereign AI, because you talk about loss of control, but right now, these models are agency promoting.”
- Publisher
- Machine Learning Street Talk
“I mean, I think Dario has this nice single sentence that we're at the end of the exponential, which I really like, in the sense that essentially what's gonna happen or, like, what seems to be what all the labs are planning for is taking these models and using those models to help build the next generation of models, and then that generation of model is more capable and can help build the next next generation, and so we're gonna get this, like, we may not have that much time before we have AIs that are smarter than all of humanity combined, for example.”
- Publisher
- Machine Learning Street Talk