Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“When I think about like how much the weights change with fine tuning, I usually think of that as kind of more a function of like some sort of divergence penalty, some sort of tethering of the, you know, the model as it's evolving to the base, to the starting point.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…we're allowing ourselves to expand to the best frontier models. At that point, I suspect for a creative writing task, you would end up in a position where you're better off using the frontier models. And yeah, we can sort of get into-- there are definitely tasks where I would say the exact opposite and say that the RL could do well. I would also say that this is obviously all dependent on the amount of compute. I think, theoretically, anything's possible. If you buy yourself a data center and spend a couple billion dollars on this task, you would be able to surpass the frontier. But the trade-off point would be fairly long, I suspect, along that curve for a task of this shape. I'd like to understand this grooves thing better, and I mean, I do-- know what you're gesturing at. When I think about like how much the weights change with fine tuning, I usually think of that as kind of more a function of like some sort of divergence penalty, some sort of tethering of the, you know, the model as it's evolving to the base, to the starting point. I think you can do that on any kind of fine tuning, right? So how is it that if I have kind of a similar divergence penalty term to my loss function, why does the supervised, why is it more destructive than the reinforcement learning? Yeah, no, that's a totally fair question. So yeah, I think what you're talking about is that there's a term called a KL defergence penalty, which is a sort of auxiliary term you can add to any loss function saying, hey, prevent the, it doesn't actually prevent the model weights from drifting. What it prevents is specifically the log probs that are generated at each token position from drifting too far from the base model. And this is often considered best practice because, you know, it can help you from getting, you know, like catastrophic forgetting and like moving too far away. However, what it's not The fundamental issue is that, let me put it this way, there are often different ways to get to the right answer, right? And the easiest example here is if you're talking about a reasoning trace where it's like, hey, you're doing a math problem and you're training this model with RL to solve the math problem. And there's probably an infinite number of ways you could reason through from a problem description to the answer. And some of them are going to be paths that the model already is comfortable with and is like, oh, these eight tokens in a row, even the model you're starting from would have generated them anyway. And then the next token, yeah, maybe it would have gotten that one wrong. And so there, the learning signal is teaching you to move that one slightly. But fundamentally, what RL just structurally optimizes for is Changing the fewest tokens the fewest log probs necessary to get to that right answer Whereas what SFT does if you're if you're doing SFT on you know You're say distilling a larger reasoning model into a smaller reasoning model And this is particularly true if like the smaller model you're going on had different pre-training distribution So so you know you would expect that kind of like it's it's kind of like built-in You know intuitions or inclinations are are different is you're not respecting those pieces of the reasoning that it would've gotten right anyway. You're overwriting the whole thing with the reasoning from the larger model. And by overwriting the entire thing, this is quite confusing potentially for the backprop algorithm because the backprop is just seeing, Oh, all of these tokens need to change, and maybe some of them didn't actually need to change. Maybe the direction the model would've gone with this token actually was also fine, but you're changing all the weights to get to this new one, and really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.