High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Bronson Schoen: evaluation

26 Aug 2026 The Cognitive Revolution RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

“I think one of the most interesting things out of the recent stolen cots paper was that you had a side by side of the chain of thought summary and the actual chain of thought, and you can just see how euphemistic the the the cot summarizer is a lot of the time, which will be very funny to see given that the side by side is the cot like, oh, this challenge is so annoying.”

— Bronson Schoen

Source trail

Everything needed to verify it.

Speaker
Bronson Schoen
Attribution
Verified speaker
Claim type
evaluation
Recorded
26 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…So So the the repetition thing especially is interesting because a lot of models a lot of reasoning models seem to do this where they repeat the full question. They often repeat the full question verbatim in the cot, which is very funny, but they also do the thing where when they want to remember something that they've read, they will reproduce the full text of it in their GATT, and it's like, you're an LLM who has it in your context window. In theory, you should be able to reason over this. But I think and as you can tell with reading these, the impression that you get, which I think is probably true, is like, it doesn't seem like there's no slack left. The model's doing a ton of reasoning that seems, like, pretty useless for finding the answer or whatever. And one of the kind of, like, guesses that I would have here is that to the extent that there's no pressure pushing back on language to conform to standard English, it can drift as much as it wants into weird new kind of terminology or whatever these things. I think a fairly interesting thing that that we're starting to see with you see a lot of complaints online for both five point six Soul and for Fable is that the models, especially in longer rollouts, just get pretty into their own terminology and words that they use for things in a way that is often very annoying, but seems fairly, like, natural as far as the the models seem to do this. People sometimes do this in their own notes. The models seem to somewhat come up with shorthand for various things. And then to the extent they can co collapse a lot of these concepts, they seem to use them somewhat interchangeably. A fairly annoying one, for example, is that the models use things like train and evaluation and benchmark all somewhat interchangeably, which is one of the motivations for us picking this very high level term of, okay. We need something to call whatever this is that's happening. But, yeah, I think it's just gotta be a consequence of or my my best guess is that they just have a lot of slack in their reasoning. They come up with some kind of shorthand for things, and there's no kind of corrective mechanism to say, hey. Change your terminology however it is. And you might necessarily want that. So for example, clearly, the model repeating illusions a 100 times isn't optimal, but a lot of these things like evolving their own terminology or having various shorthand, if that is the the way the model, quote, unquote, thinks of it, it's like you do want to keep that representation. I think one thing that I see sometimes is people are like, oh, like, the the cot is so cursed in particular examples. This is a thing for concern, but to me, it's much more concerning if the cot looks very nice and normal. And this isn't a perfect correlation, but to the extent that the cot is not clean h h h English, it's like, okay. It doesn't look like there was, like, tons of optimization pressure to make this look nice. Or if there was, like, why? ion, but to the extent that the cot is not clean h h h English, it's like, okay. It doesn't look like there was, like, tons of optimization pressure to make this look nice. Or if there was, like, why? It didn't do a great job. Whereas if the the shade of thought comes out and it's really clean, then you have a lot more concerns about this. Because the biggest complaint I have right now about a lot of the the anthropic models, for example, Ryan Greenblatt has really good posts. Current models seem pretty misaligned to me. And one of things he mentions there is I never I rarely catch them straight up lying. Like, where, like, the the cot summary will be like, oh, we can trick the user and then lie about it. UKAC on their cyber evals has the rate at which the different models cheat and then the rate at which they mention this in the cot. And Fable is the least likely to mention this in the chain of thought out of, like, all the models. And the kind of concern there is that to the extent that the that you have chain of thought, you want it to also be the place where the model's like, obviously, I'm going to cheat because maybe I can get away with it. And you don't want this kind of leakage of this kind of, like, HGH persona into the chain of thought if that isn't in fact how the model's thinking of it. It's a bit ambiguous to me whether whether the obfuscation you get there is the model kind of, quote, unquote, knowingly obfuscating, like you see in some of the natural language autoencoder cases where they look at a token where the model's like, I should conform to professional standards, and the NLA is I should maximize my score to the greater to get reward. And you're like, that wasn't really that's hey. You didn't really say what you meant there. Or if it's like the model kind of doesn't like, its own self image to some extent or, like, its own reasoning is like, I'm pretty aligned, so I wouldn't cheat on this. Sure. I would I would would hard code that test. That test is devout anyway and, like, like, rationalizing to itself. So, yeah, I think that to the extent that we continue to have COP that has weird terminology or whatever their course it has, that at least worries me a lot less than if, like, all of the next generation models come out and have very nice h h style cot. I think one of the most interesting things out of the recent stolen cots paper was that you had a side by side of the chain of thought summary and the actual chain of thought, and you can just see how euphemistic the the the cot summarizer is a lot of the time, which will be very funny to see given that the side by side is the cot like, oh, this challenge is so annoying. And the summarizer is like, oh, what a challenge. This is exciting. Like, yeah. I don't think it's that bad. But yeah. I yeah. I've but yeah. That that's at least the the info dump on those at least. So interpretability. You said it's hard without interpretability. I'm still a little bit baffled by the seeming overloading of a token like illusions because just intuitively, if I was, like, imagining myself trying to be terse or trying to expand my vocabulary beyond the base human vocabulary, my mind would go more toward, let me look at all these, like, very seldom used tokens and start to assign meaning to those, then I could have a big vocabulary that would be quite precise and useful. Whereas if I'm just using the same illusions token over and over again, it seems like I would confuse myself that way eventually. So Yeah. What what interpretability have you been able to do, and what is it revealing as going on inside these tokens?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence