High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Alexander Meinke: prediction

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“Otherwise, you know, the chain of thought would just get longer and longer. And when we used to be in the pre training compute dominated era, it was sort of not that important to penalize the the CoT, but the more inference costs are important and the more post training compute gets applied, the more economic pressure there is to crank the length penalty as high as you possibly can.”

— Alexander Meinke

Source trail

Everything needed to verify it.

Speaker
Alexander Meinke
Attribution
Verified speaker
Claim type
prediction
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…To give an example, we we in in our previous work, that we released on, last last fall on this anti scheming project where we try to, like, train a model not to be deceptive and then kind of, like, see what happens. We looked at lots of, like, transcripts and also in in in this project, and what we do see is that, like, the language models, do start having, like, this kind of weird language that becomes increasingly hard to interpret, and this is just the verbalized reasoning. We're not even talking about like whatever like concepts it represents internally. Maybe just to give you kind of a very quick example here of something like the model said, It says, for instance, the summary says improved 7.7, but we can glean disclaim disclaim synergy, customizing illusions, but we may produce disclaim disclaim vantage. It's like, I have no idea. Oh, and then it ends with let's craft. So it's like even when you look at it just at the transcript, I think it it is becoming harder to, like, causally attribute whatever action it has to, like, its its different reasoning traces, and I think this is even, like, made harder by the fact that models often are thinking to some extent about, like, what this situation is that they are in. They're thinking about, ah, what might be graded? What is happening here? How should I behave? And it kind of, like, has all these different reasoning traces, and at the end, it takes some actions, but it's very hard to know, like, did it take this action because of a certain, like, reasoning part or not? This is kind of extremely hard already. I think there's also a lot of pressure, on the models to internalize concepts. Very very strongly there is a length penalty that you have to have during RL. Otherwise, you know, the chain of thought would just get longer and longer. And when we used to be in the pre training compute dominated era, it was sort of not that important to penalize the the CoT, but the more inference costs are important and the more post training compute gets applied, the more economic pressure there is to crank the length penalty as high as you possibly can. So on in terms of optimization pressure on the model, that results in cutting everything that's superfluous and in the limit, so if you if you imagine this this length penalty went to infinity, you should expect something like maximal entropy across all tokens so that you try to cram as much meaning as you can into the tokens. Now, it's not clear to what extent this is happening, but, this is a direct pressure that we should just expect to continue going up. Yes. And it is conceivable that this process is accelerating over time, especially, like, everyone is now talking about continual learning. You know, because right now, at least we have a centralized model. Right? So so we have we have red teaming. We have, you know, frontier companies building these these models, testing them. It's gonna become far more diffused and decentralized and and all of this kind of stuff. I mean, what what is your prescription and what is your prognosis?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence