Evidence receipt / evaluation
Published · transcript-backedKyle Corbitt: evaluation
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“Like one open question was sort of like, hey, how do we do length normalization? And this comes into if you have a trace that happens to be, so, so it's sort of the original math in GRPO actually structurally advantaged very long thinking traces and, and, and, you know, just like generations in general, just because, you know, it didn't normalize by the number of tokens.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, okay. Is it worth getting into what some of the finer points have been since? GRPO that has made it even better, not in a maybe super mathy way, but, you know, like what additional insights have people brought to bear since then? Yeah, I mean, yeah, we can talk about it briefly. Yeah, it's a bunch of small things. Like one open question was sort of like, hey, how do we do length normalization? And this comes into if you have a trace that happens to be, so, so it's sort of the original math in GRPO actually structurally advantaged very long thinking traces and, and, and, you know, just like generations in general, just because, you know, it didn't normalize by the number of tokens. So basically, if you had a batch of, you know, like, whatever, like 128 different completions, and one of the traces happened to five times as long as the others, it ended up with like five times the amount of weight in the way the models were updated than others. And so it's sort of like, you know, the people had pretty good success with basically kind of like down weighting that to average it out. You know, there were You know, CISPO's a really cool one. It basically just changes the way you're doing the clipping. PPO and the GRPO inherits this, has a specific way of making sure that the weights don't stray too far in any one round of updates. And there was this new technique called CISPO that was released maybe six months later or something that basically it puts the clipping in a different spot, which basically like... The idea there is it lets the model discover much more quickly those like very, very high value but rare tokens. And so it sort of allows those to update the weights much more if there's a very high score while not like allowing the weights to update too much. So yeah, and then there's like a stack of probably like, I don't know, half a dozen kind of like little tricks like that that people have developed to make the algorithm both more stable and converge faster. Amazing. That's been a great trip down the rabbit hole. Popping out now again and trying to think about what it all means. Obviously, the huge thing about reinforcement learning that we've seen time and time again, but you know, it's really happening now. I'm just thinking about this latest Erdos problem that's been solved in the last 24 hours, or at least reported, that Rune just said something like, This is the first time that everybody in the math community is super impressed. And the key point that I'm getting at here is reinforcement learning has the ability to take a model past what available training data has on offer to teach it, right? So this is where we get superhuman performance. Now, how does that happen? I mean, you had kind of talked about the grooves and by focusing in on these key decision points rather than just mashing every token, you're kind of playing to the model's established strengths. But clearly there's like also something happening where the at scale, the reinforcement learning is teaching qualitatively new capabilities to the model. So how should I think about that? You know, in other words, how are we How are we, where clearly it's, well, I mean, you could argue with me if you think this is wrong, but I take it that everybody kind of has come to accept that this is where the superhuman performance comes from. But I don't have a great intuition for where we're making that move from playing to the model's strength, staying in the groove, you know, focusing on what matters and reinforcing what it already knows or has at least some instinct for into this like qualitatively new regime where now we're solving open math problems.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.