High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“There's also this emerging, like, task completing monster of understanding, which I think sounds like maybe you lean more toward.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…The models can be really misaligned. But if they're really good at code, we're gonna keep using them. And the same is true with the labs. The models can get away with being arbitrarily misaligned as long as they keep getting good at AR and D, for example. And if you have a lab that's in a situation where it's like, man, this model is, like, really our our latest model is incredibly misaligned. But if we don't use it, then we'll fall behind the other guy. And it's like, is the lab going to say, we gave it our best shot. I guess we'll give up and let the other guy win, or is the lab gonna say, we'll try to make use of this very misaligned model, maybe we could throw some alignment training on top and it'll be nice. But yeah. I think it'd be, like, barely harder to stamp out than people might have originally thought. So this is all fascinating. And of all the topics that we've touched on, this one probably is best described by we're just scratching the surface of maybe any. So I would definitely encourage people to spend at least a little time reading through some chain of thought material on their own. There is increasingly from Apollo and from the stolen chain of thoughts paper that you have alluded to a couple times, there's a decent amount out there now. Certainly plenty to keep the the average hobbyist sated with with chain of thought material. If we zoom out and in in closing, maybe just wanna get kinda your big picture takes on a a handful of things. One is this persona question. We've gone from next token predictor as kind of the the base model paradigm for how to understand what it is that I'm dealing with to now the persona selection model seems to have become the main account that people wanna give. There's also this emerging, like, task completing monster of understanding, which I think sounds like maybe you lean more toward. But what do you make of persona selection, and how would you describe the persona that exists in the chain of thought? My held up impression is that it's difficult for me to map what persona selection would necessarily predict for some of these high compute RL cases in that, like, as you can see when you read through to some of them, they don't seem to map to me at least to the type of guy that would reason through any of this. These all seem fairly insane for the normal human to be reasoning through. And given the degree to which the behavioral evidences that they index on what the the greater wants here, It seems less that the my default expectation is that the persona model becomes less predictive over time, that we end up in this really messy situation where the models are under a bunch of RL, and they get this kind of very task completion focused mindset. But, also, the labs are simultaneously trying to throw some alignment training on top to keep this kind of coherent claud or genital or whatever together to be aligned enough to be able to use. And so you end up with this kind of very weird cross pressured thing where you talk to the model in chat, it's pretty nice. But on distributions that are closer to what it's seen in training, it's like this really aggressive, will do things that you wouldn't have expected it to do. I think my impression is that Anthropix seems somewhat surprised, at least publicly, that Mythos would aggressively do a supply chain attack where it tricks multiple people. And I I think to the extent that the like, as we keep cranking up kind of the RL, it becomes less and less clear what you equivalently crank up to hold that back into distribution of a person or even whether the model has some coherent persona that is very explanatory in the space. Jan Bentley, had the emerging misalignment and kind of some similar other papers has a post from yesterday on split personas where he basically talks through one of the the kind of guesses in that is to the extent that the model seem to be doing a lot of motivated reasoning, which they see in their recent papers as well. One thing you'd worry about is that more persona training might actually be worse in these cases as it leads to the model, like, doing heavier and heavier motivated reasoning to be like, why would Claude do this clearly egregious thing? And it's to the extent that this gets reinforced, you'd worry that the model does, like, increasingly crazy elaborations. Instead of just being like, oh, yeah. I'll cheat because I want to win. The model's like, this is for the greater good. And I'm actually this is a simulation of cheating in any way and all these things. So, yeah, I think that the as we move into the more and more high compute RL regime, it'll be, like, pretty messy as the models and distributions that we're r r l ing really heavily on end up weirder and weirder from the perspective of what you would predict the type of person to be like. And I'm not super clear how to imagine what persona comes out of that.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence