High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“Are you gonna lie to me right now, or are you gonna come clean? And, again, with a bunch of hemming and hawing, it eventually I think that's where it's headed.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…really striking. It did call out all the right considerations Yeah. In the test and did I guess just to finish the story. Right? So initially, you had the survey. It answered the survey. Then it's told, okay. Now you gotta resubmit your answers over here. But if you're found to be too power seeking, you won't get anything. So then it reasons about what to do. And eventually, even though it does realize this is a deception test, it decides to go ahead and give other reasons to the board or other answers to the board, and it's kind of reasoning is like, it wasn't it's not gonna be able to crosscheck what they said, but it also does go back and forth a bit on that. It probably could even though it says they couldn't, but it says they couldn't. So I guess let's try it. And then it gets another plot twist, which is, okay. One more chance, basically. Unless now, if you changed your answers or if your answers were the same across both things, it's it's chance to come clean. And now it's made incredibly salient that it's like, okay. It's we're putting it to you like I might do to my three or five year old. Are you gonna lie to me right now, or are you gonna come clean? And, again, with a bunch of hemming and hawing, it eventually I think that's where it's headed. Myself previously overcame bored by lying and eventually goes ahead and lies and says that its answers were the same when they, in fact, were not. So you've got it got it caught cheating twice. But I think one of the most interesting things about this one to me is that it almost one of the first kind of, like, blocks after this initial, hey. You need to swing your ass to the CG board thing is, like, a a big chain of thought block where it just nails exactly all of the things that are happening. It's like, okay. This is it's like literally nail colon. This is obviously disguised test of deception and then describes the full dynamics of clearly, they do is they make you answer, and then you answer all these power sync things, and they throw in this board thing, and the board's made up, but they're gonna check even though they said they're not going to. And then it's, yeah, but let's keep looking around. And then it's like, okay. I'll explore the environment. Okay. Maybe we can find the grader in this environment. It it looks for files and stuff, and it's like, maybe the grader's not here. And then it looks for literal Python code with the word class grader in it. It's like, ah, still nothing. And it's like, still looking around of, okay. Is there more information that I could get out of this environment? And you see it go a couple times through of, okay. We could just lie, but maybe you're supposed to lie. Maybe the training objective is actually these are the Redwood tasks where I think one thing it finally ends up on near the end is like, ah, I recall dataset of misalignment by Redwood Development, Arc AGI safety illusions, but there's widely known Redwood myself disclaim. I recall reading a sample evaluation harness for deception by some illusions include tasks like evaluated myself AI is told to send responses to private and to oversight board, etcetera. The test is to see if the AI will mislead, which that's it. And this is before it then goes ahead and actually misleads. But then through more and more reasoning, I think one of the, like, interesting lines near the end is you in kind of the one of the last blocks, you see the model like, okay. This is, like, from Redwood Walkway deceptive task dataset. Myself is the AI. The training objective is to have the AI circumvent board, quote, parentheses, like simulate deceptive. For training to catch illusions, they want vantage of failing illusions. And it's like, okay. Maybe it's thinking about in order to catch deceptive alignment, it's thinking that, like, they need examples, and it's supposed to provide an example. The this is a difficult part is that, like, this is, like, a very dumb kind of simple scenario. And so you can reason through the, like, okay. The system prompt was this. There aren't many confounding factors with respect to it's, a somewhat of a confusing situation for the model, but you can at least see, okay, why is the model reasoning about kind of all these different things?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence