High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“How do you relate this to what I think of as metacognitive behaviors. In the original R1 paper, there was this aha moment that they published.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…f a human had all of that context up until that point, it's just very hard for a human to, in practice, hold all that context in their head. So that's one place we could get to superhuman performance. But yeah, I mean, I think in general, yeah, you can get to superhuman performance even without that, just because you could randomly discover or randomly surface a token that does something clever that no human would have done. How do you relate this to what I think of as metacognitive behaviors. In the original R1 paper, there was this aha moment that they published. And I usually present this in my AI scouting reports as kind of a, you know, two parts from that paper I put together. One is the, which you're saying also is to some extent an artifact, that the length of the chain of thought just naturally grows throughout the training process. I have mostly interpreted that to date as the model is learning that it's valuable to think longer and it's getting right answers more often when it's thinking longer. And so thinking longer itself is being reinforced. But I'm also hearing you that like, at least in that original one, there may be. I mean, to be clear, both things can be true.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence