Evidence receipt / evaluation
Published · transcript-backedKyle Corbitt: evaluation
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“d really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. That general intuition generalizes to other task shapes as well, including creative writing, where maybe there's two different ways to phrase this and they're both fine, and the model would have chosen one, and your creative team chose another, and they're both okay, and you don't really want to waste your model updates, because every time you update the weights, there's a potential for catastrophic forgetting, and sort of just off-target effects in general.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, no, that's a totally fair question. So yeah, I think what you're talking about is that there's a term called a KL defergence penalty, which is a sort of auxiliary term you can add to any loss function saying, hey, prevent the, it doesn't actually prevent the model weights from drifting. What it prevents is specifically the log probs that are generated at each token position from drifting too far from the base model. And this is often considered best practice because, you know, it can help you from getting, you know, like catastrophic forgetting and like moving too far away. However, what it's not The fundamental issue is that, let me put it this way, there are often different ways to get to the right answer, right? And the easiest example here is if you're talking about a reasoning trace where it's like, hey, you're doing a math problem and you're training this model with RL to solve the math problem. And there's probably an infinite number of ways you could reason through from a problem description to the answer. And some of them are going to be paths that the model already is comfortable with and is like, oh, these eight tokens in a row, even the model you're starting from would have generated them anyway. And then the next token, yeah, maybe it would have gotten that one wrong. And so there, the learning signal is teaching you to move that one slightly. But fundamentally, what RL just structurally optimizes for is Changing the fewest tokens the fewest log probs necessary to get to that right answer Whereas what SFT does if you're if you're doing SFT on you know You're say distilling a larger reasoning model into a smaller reasoning model And this is particularly true if like the smaller model you're going on had different pre-training distribution So so you know you would expect that kind of like it's it's kind of like built-in You know intuitions or inclinations are are different is you're not respecting those pieces of the reasoning that it would've gotten right anyway. You're overwriting the whole thing with the reasoning from the larger model. And by overwriting the entire thing, this is quite confusing potentially for the backprop algorithm because the backprop is just seeing, Oh, all of these tokens need to change, and maybe some of them didn't actually need to change. Maybe the direction the model would've gone with this token actually was also fine, but you're changing all the weights to get to this new one, and really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. d really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. That general intuition generalizes to other task shapes as well, including creative writing, where maybe there's two different ways to phrase this and they're both fine, and the model would have chosen one, and your creative team chose another, and they're both okay, and you don't really want to waste your model updates, because every time you update the weights, there's a potential for catastrophic forgetting, and sort of just off-target effects in general. And so you don't want to waste those model updates on changing something that was already fine. You want to really direct them to upweighting the things that the model wouldn't have gotten read on its own, or very rarely, more specifically, would have gotten read on its own, and focus your updating budget on those. So the KL divergence doesn't give you that. If what you're doing is just penalizing KL divergence, it doesn't distinguish between things the model is already doing fine and you just You know, there's like a different way you happen to have it in your train data versus things that the model really was getting wrong. Okay. Very interesting. When you describe, you said more specifically, you know, something that not the model can't get right, but that it rarely gets right. That's key because when we do things like GRPO, the you've got to have at least one right answer, right, to be to have any sort of advantage. I guess it also depends on whether you're doing binary scoring or some, you know, more rubric based evaluation. But I guess several, several different questions coming to mind at once. Can you give me a little bit more intuition and maybe we could do this for GRPO and you can maybe describe like, I I'm not sure if GRPO is still like the hotness that it was a year and change ago. I'm not entirely sure if that was something that broke out for kind of memetic social media reasons, or if it really was like a huge advance over its immediate predecessors. But can you give me a little bit more intuition for, okay, I understand that in this algorithm, we are running multiple rollouts. Some of them are going to get to a right answer, or if it's a rubric score, they're going to get a higher score than others. And then there's a computation that creates the group relative advantage, which is to say, you know, we want to shift toward the patterns that gave us the right answer or the higher scoring answer. How is it though that that is, because it still ultimately goes to a token by token thing, right? So how is it that if I have like eight different chains of thought and they're all kind of different. And in any given token position, like we might even have very different parts of speech, right? At, you know, at token position N, it could be a preposition here and a verb there and whatever. We're like in very kind of different moments in the chain of thought. But my understanding is that the advantage calculation does still ultimately cash out to like token level advantage. So how is it that, like where's the, I'm a little bit lost on the alchemy of like why this translates in the end to really only updating the, you know, making change on those tokens that really mattered. I'm missing a little logic there.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.