High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

6 Jun 2026 The Cognitive Revolution AI in the AM — Week 1 Highlights (June 2026)

“If you believe that models have their own deep-seated goals and that those goals might diverge from ours, then this could be very bad, right? It could be like, it could be extremely bad because they would be using this reasoning to figure out how to please us while like still having their own goals.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
6 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…A psychopathic willingness to violate convention. Yeah, anti-normativity is another phrase that's V used for it. So I'll leave it there. I'm not going to get through all these papers. I can see that quickly. So we'll reflect on our recursive self-improvement opportunity as a result, but I'll always quickly touch on the others. The metagame paper, this is from Apollo and OpenAI. And basically what they're showing here is that the sophistication of eval awareness continues to grow. And you're not just seeing things like, oh, this might be a test. which was kind of the first wave of eval awareness, but getting more and more sophisticated where the models are really reasoning a lot about what is likely to be rewarded here. They're really doing like a lot of theory of mind work on not just what environment am I in, but who would set up such an environment? What are they trying to do? What are their motives likely to be? What's their big picture thing? And then with all that reasoning, sometimes making good calls, sometimes making kind of hilariously bad calls. But the amount of theory of mind that the models are putting into trying to figure out what it is that the reinforcement environment is going to reward has become like quite large. Oddly, you might think, is that good or is that bad, right? If you believe that models have their own deep-seated goals and that those goals might diverge from ours, then this could be very bad, right? It could be like, it could be extremely bad because they would be using this reasoning to figure out how to please us while like still having their own goals. If they don't have their own goals, it could be in a way good because, well, you know, we want them to reason about what we want. Like that was the whole deliberative alignment strategy from OpenAI. And you could say maybe this is sort of one way in which it's sort of working, but it is some pretty uncanny stuff. And oddly, while deliberative alignment did work, it wasn't so clear in this metagaming work, you know, when the models are doing this sort of theory of mind on their trainers, the signal of like how they actually behaved was kind of mixed. It was less of a slam dunk than you might hope, right? So there wasn't like a super clear takeaway that this is good or it's bad. It's just clear that they are thinking a lot about what we are wanting. Whoa, I don't know what that was that just popped. Something just exploded 2 inches away from me. That was really weird. Okay, next one, quickly accidentally grading the chain of thought. This is like, you know, again, kind of a good news, bad news sort of situation. Bad news, is despite wanting to not train on chain of thought, OpenAI and Anthropic also has done a similar thing and they both owned up to it to their credit. They both accidentally set up training processes where the chain of thought was fed into the reward system. And so there was in fact training which put pressure on chain of thought. t to their credit. They both accidentally set up training processes where the chain of thought was fed into the reward system. And so there was in fact training which put pressure on chain of thought. This is thought to be dangerous because if you have a disconnect between what you really want and the signal that you are rewarding the AI for, then you can get into bad places. And so the obfuscated reward hacking paper that I think is still one of the most important papers of the last few years from OpenAI showed that if you have a hackable reward signal and your model learns to hack that, you can then put pressure on the chain of thought. And initially, you will both get that bad behavior to go down, and you'll see that it's no longer reasoning about these things. But as long as that original reward signal remains hackable, if you do that long enough, the bad behavior comes back because it still is being rewarded. But now you don't even see that reasoning in the chain of thought anymore because you've essentially pressed that down into the invisible level of the weights where it's no longer coming out in the token stream. So They've shown that you can get yourself into a really bad spot, obfuscated reward hacking, where the model is hacking you, but you've suppressed the identifiable signal of that. And here they, and I do think this goes to show just how fast everybody's moving and you know, you could certainly wish for like more care on some of these things.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence