High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Nathan Labenz: recommendation

6 Jun 2026 The Cognitive Revolution AI in the AM — Week 1 Highlights (June 2026)

“This is thought to be dangerous because if you have a disconnect between what you really want and the signal that you are rewarding the AI for, then you can get into bad places. And so the obfuscated reward hacking paper that I think is still one of the most important papers of the last few years from OpenAI showed that if you have a hackable reward signal and your model learns to hack that, you can then put pressure on the chain of thought.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
recommendation
Recorded
6 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, anti-normativity is another phrase that's V used for it. So I'll leave it there. I'm not going to get through all these papers. I can see that quickly. So we'll reflect on our recursive self-improvement opportunity as a result, but I'll always quickly touch on the others. The metagame paper, this is from Apollo and OpenAI. And basically what they're showing here is that the sophistication of eval awareness continues to grow. And you're not just seeing things like, oh, this might be a test. which was kind of the first wave of eval awareness, but getting more and more sophisticated where the models are really reasoning a lot about what is likely to be rewarded here. They're really doing like a lot of theory of mind work on not just what environment am I in, but who would set up such an environment? What are they trying to do? What are their motives likely to be? What's their big picture thing? And then with all that reasoning, sometimes making good calls, sometimes making kind of hilariously bad calls. But the amount of theory of mind that the models are putting into trying to figure out what it is that the reinforcement environment is going to reward has become like quite large. Oddly, you might think, is that good or is that bad, right? If you believe that models have their own deep-seated goals and that those goals might diverge from ours, then this could be very bad, right? It could be like, it could be extremely bad because they would be using this reasoning to figure out how to please us while like still having their own goals. If they don't have their own goals, it could be in a way good because, well, you know, we want them to reason about what we want. Like that was the whole deliberative alignment strategy from OpenAI. And you could say maybe this is sort of one way in which it's sort of working, but it is some pretty uncanny stuff. And oddly, while deliberative alignment did work, it wasn't so clear in this metagaming work, you know, when the models are doing this sort of theory of mind on their trainers, the signal of like how they actually behaved was kind of mixed. It was less of a slam dunk than you might hope, right? So there wasn't like a super clear takeaway that this is good or it's bad. It's just clear that they are thinking a lot about what we are wanting. Whoa, I don't know what that was that just popped. Something just exploded 2 inches away from me. That was really weird. Okay, next one, quickly accidentally grading the chain of thought. This is like, you know, again, kind of a good news, bad news sort of situation. Bad news, is despite wanting to not train on chain of thought, OpenAI and Anthropic also has done a similar thing and they both owned up to it to their credit. They both accidentally set up training processes where the chain of thought was fed into the reward system. And so there was in fact training which put pressure on chain of thought. t to their credit. They both accidentally set up training processes where the chain of thought was fed into the reward system. And so there was in fact training which put pressure on chain of thought. This is thought to be dangerous because if you have a disconnect between what you really want and the signal that you are rewarding the AI for, then you can get into bad places. And so the obfuscated reward hacking paper that I think is still one of the most important papers of the last few years from OpenAI showed that if you have a hackable reward signal and your model learns to hack that, you can then put pressure on the chain of thought. And initially, you will both get that bad behavior to go down, and you'll see that it's no longer reasoning about these things. But as long as that original reward signal remains hackable, if you do that long enough, the bad behavior comes back because it still is being rewarded. But now you don't even see that reasoning in the chain of thought anymore because you've essentially pressed that down into the invisible level of the weights where it's no longer coming out in the token stream. So They've shown that you can get yourself into a really bad spot, obfuscated reward hacking, where the model is hacking you, but you've suppressed the identifiable signal of that. And here they, and I do think this goes to show just how fast everybody's moving and you know, you could certainly wish for like more care on some of these things. sed the identifiable signal of that. And here they, and I do think this goes to show just how fast everybody's moving and you know, you could certainly wish for like more care on some of these things. They did it by mistake. not a huge portion of the data, but kind of low single digits for different models, it varies, were trained this way. And basically what they found is there's at least some tolerance for mistakes. You know, this did not create a very bad result in the models that were trained this way. So that's sort of good. You know, it's one example where we might think physics is sort of being kind to us. Like it's not, if you just do a little bit of it, you know, you don't poison the whole well. I would say there's still some caveats there. Do we really know that there's no issue? No, we just know that this investigation didn't find flagrant issues. I also do worry a little bit that it will lead people to be a little more careless than they otherwise would be. This was supposed to be a strong taboo. We violated it. Now we're like, oh, well, maybe it wasn't so bad that we violated it. What's that going to do for the power of the taboo in the future? And then what's the solution to this? The solution is We've got new automated systems, more monitoring, right? So OpenAI has now set up monitoring on top of monitoring to try to detect if the chain of thought is ever being used. And this is really emblematic of like their strategy for everything. Just if we have a problem, throw an AI monitor on it and you know, hopefully it'll catch it. And then we can go back to, you know, pushing toward recursive self-improvement as fast as possible. I'll do one second on the 4th one and then we'll skip the 5th one and we'll get to Matthew because he's here. This natural language autoencoders thing I think is really exciting. If you're worried that your model is thinking thoughts that it's not expressing in tokens and that those thoughts might be problematic, then one way you might try to get at that is do some sort of internal monitoring. Can I look at the internal states, make sense of them and detect problematic things there? There's been a lot of strategies that try to do that. Like everything else, they sort of work. They don't fully work. But a challenge is interpreting the internal states, obviously. With the natural language autoencoders, they basically set up a system where the model must pass through natural language as part of its forward pass. and using a reconstruction loss, which basically means like the model has to both kick out to natural language and then get back from natural language and still do its original task in the same way that it was always going to do it, they're now able to kind of get these little short paragraph length things that represent in natural language what the model is thinking at any given moment in its inference rollout. And then they can look at that and it is much more human readable than certainly like, here's a sparse autoencoder with like, these features lit up and these features, by the way, were maximized by these other passages in the training data. And so, we kind of squint at it and think this or that.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence