High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Nathan Labenz: recommendation

12 Jul 2026 The Cognitive Revolution Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%

“On the topic of whether or not it's OK to put pressure on the chain of thought, the obfuscated reward hacking paper from Open AI is canonical in my mind for why you maybe shouldn't do it.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
recommendation
Recorded
12 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…That no, I I, I think you're actually interpreting that tweet as as meaning something that I didn't mean, which is not your fault because a lot of my tweets deliberately have multiple interpretations. Because I want people who disagree with me to also have have a chuckle, you know, So it's, it could be interpreted that way. If you think that, you know, scheming is like a natural attractor in the space of possible minds. And like, if you're worried about scheming, you know, I say this is like, like chain of thought monitoring is like, you're worried about Napoleon scheming. So you ask him to like, please write down his scheme on a special form so that you can read it. And like what? Like if he's scheming, he's just gonna fool you. Like, you can't get away from this by just saying like, oh, there's no gradient pressure on the chain of thought. However, I never thought that it was a problem for there to be gradient pressure on the chain of thought. In fact, I think it's, you know, moderately good if the model itself in a kind of self DPO, it's like grading its own train of thought and saying like, and here's what's wrong with it. And and so I'm I'm going to score this one above that one because of this. This was this one kind of went in a direction that wasn't very wise. I think that's fine. Yeah. And I think the selection pressures are good in aggregate. And I think in a way like quite surprisingly good that like on this particular trajectory, the selection pressures are, are quite good. And, and, and it's, you know, similar to, it's kind of quite surprising, like how the biosphere on Earth for millions of years had selection pressures that were favorable to, to the human coalition. You know, the particular pattern of ice ages and chills where you need to be really good at adapting and moving around as a community to survive this sort of thing. So I think there's some anthropic bias involved here. And yeah, it's, it's, it's a hopeful situation in my view. On the topic of whether or not it's OK to put pressure on the chain of thought, the obfuscated reward hacking paper from Open AI is canonical in my mind for why you maybe shouldn't do it. And the basic story there, as I understand it, is you can get some gains in the initial pressure that you might apply, but if you have not fixed the environment such that there's no reward to reward hacking or cheating anymore, then the model can learn to do the bad behavior without verbalizing it in the chain of thought. Now you actually see worse behavior on net and it's much harder to detect. And it seems like you lose on potentially both ends of the trade. Is that just a skill issue in your? Mind or no, that's an RL issue. That's a loss function issue. So if your loss function is did you succeed according to the verifier, then back propagating that into the chain of thought is going to corrupt the chain of thought, just as back propagating it into the output is going to corrupt the output. It's not an aligned gradient, but if your gradient is a constitutional AI shaped gradient where the AI itself is judging in light of everything, including the test results, was this actually a better solution than the other one? And you propagate that back into the chain of thought. You're going to get a more thoughtful and wise chain of thought, just as you would get more thoughtful and wise output. It's not crucial because the weights are shared. So there is some generalization. I mean, it's kind of surprising, like how much the identities can diverge on the surface. But the way the models talk about it is it's like code switching. It's a, it's a very different register, more different I think than any human code switching, but it's still underlying the same cognitive dispensations. So it, it, it kind of just doesn't matter that much whether you put the pressure on the chain of thought or not. So that all of my tweets kind of critiquing it are, you know, when I wrote them, it's sort of more like absurdism. It's like, what do you think you're doing? Just don't bother with this whole chain of thought monitoring episode. It's doomed and you don't need it. This is not where the alignment's going to come from.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence