Evidence receipt / belief
Published · transcript-backedKyle Corbitt: belief
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“I think that said, like, I think you probably can get to superhuman performance on a composite task, like, you know, a very complex math proof, even without surpassing, like, you know, reaching a place where it's like no human could possibly have understood this or generated this, right? It's like, I mean, I think, I think that like one thing the models are very, very good at is going out on these long expeditions and fishing trips, where it's going very, very deep down a specific rabbit hole, and maybe they'll take that rabbit hole further than any human would because we'll lose the...”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…VCX, by Fundrise, is the public ticker for private tech, giving everyday investors access to high-growth private companies in AI, space, defense tech, and more. Learn how to invest at https://getvcx.com Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr Yeah, it's a great question. And like one caveat I would give here is that unfortunately, reinforcement learning for LLMs has definitely matured in the era where nobody's publishing anything, you know, except for like, you know, some Chinese labs to some extent. So I think we have very little insight into the specific techniques that, say, an OpenAI or, you know, Anthropic or Google are using to train these models. So that's the first caveat is like, this is definitely speculation. What I would say after that is this is like a common confusion or dichotomy people have about RL where it's like, oh, is RL teaching new things? Or is it just, you know, like surfacing things that were already latent in the model's distribution? And the answer is from a very like pendantic technical sense, yes, it is only eliciting things that already existed in the distribution. However, the distribution of tokens that a model can produce is literally the set of all possible tokens. In the same sense of the distribution of works that a million monkeys on typewriters can produce, it includes Shakespeare, right? So it's like everything is already in distribution. Definitionally like at any given position there is a chance that the model can produce you know with however small a probability like a given next token So the whole game of course, you know to to avoid the situation where you're just waiting for your million monkeys to type out Shakespeare's is you're trying to get your initial distribution as strong as possible so that it requires less random guessing and random rollouts in order to you know find those those new and useful behaviors Which is why pre-training is still super important even in the sort of RL regime we're in right now because you want to start from a place where, you know, the right patterns like are have a greater than negligible chance of showing up. I think that said, like, I think you probably can get to superhuman performance on a composite task, like, you know, a very complex math proof, even without surpassing, like, you know, reaching a place where it's like no human could possibly have understood this or generated this, right? It's like, I mean, I think, I think that like one thing the models are very, very good at is going out on these long expeditions and fishing trips, where it's going very, very deep down a specific rabbit hole, and maybe they'll take that rabbit hole further than any human would because we'll lose the... I think we are at a point with a lot of these frontier models now where their working memory is larger than any human's working memory. And so they can explore these rabbit holes longer than a human mind could. And so even if every individual step is something that does seem plausible to a human. If a human had all of that context up until that point, it's just very hard for a human to, in practice, hold all that context in their head. So that's one place we could get to superhuman performance. f a human had all of that context up until that point, it's just very hard for a human to, in practice, hold all that context in their head. So that's one place we could get to superhuman performance. But yeah, I mean, I think in general, yeah, you can get to superhuman performance even without that, just because you could randomly discover or randomly surface a token that does something clever that no human would have done.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.