Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“When you describe, you said more specifically, you know, something that not the model can't get right, but that it rarely gets right. That's key because when we do things like GRPO, the you've got to have at least one right answer, right, to be to have any sort of advantage.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…d really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. That general intuition generalizes to other task shapes as well, including creative writing, where maybe there's two different ways to phrase this and they're both fine, and the model would have chosen one, and your creative team chose another, and they're both okay, and you don't really want to waste your model updates, because every time you update the weights, there's a potential for catastrophic forgetting, and sort of just off-target effects in general. And so you don't want to waste those model updates on changing something that was already fine. You want to really direct them to upweighting the things that the model wouldn't have gotten read on its own, or very rarely, more specifically, would have gotten read on its own, and focus your updating budget on those. So the KL divergence doesn't give you that. If what you're doing is just penalizing KL divergence, it doesn't distinguish between things the model is already doing fine and you just You know, there's like a different way you happen to have it in your train data versus things that the model really was getting wrong. Okay. Very interesting. When you describe, you said more specifically, you know, something that not the model can't get right, but that it rarely gets right. That's key because when we do things like GRPO, the you've got to have at least one right answer, right, to be to have any sort of advantage. I guess it also depends on whether you're doing binary scoring or some, you know, more rubric based evaluation. But I guess several, several different questions coming to mind at once. Can you give me a little bit more intuition and maybe we could do this for GRPO and you can maybe describe like, I I'm not sure if GRPO is still like the hotness that it was a year and change ago. I'm not entirely sure if that was something that broke out for kind of memetic social media reasons, or if it really was like a huge advance over its immediate predecessors. But can you give me a little bit more intuition for, okay, I understand that in this algorithm, we are running multiple rollouts. Some of them are going to get to a right answer, or if it's a rubric score, they're going to get a higher score than others. And then there's a computation that creates the group relative advantage, which is to say, you know, we want to shift toward the patterns that gave us the right answer or the higher scoring answer. How is it though that that is, because it still ultimately goes to a token by token thing, right? So how is it that if I have like eight different chains of thought and they're all kind of different. And in any given token position, like we might even have very different parts of speech, right? At, you know, at token position N, it could be a preposition here and a verb there and whatever. We're like in very kind of different moments in the chain of thought. But my understanding is that the advantage calculation does still ultimately cash out to like token level advantage. So how is it that, like where's the, I'm a little bit lost on the alchemy of like why this translates in the end to really only updating the, you know, making change on those tokens that really mattered. I'm missing a little logic there. Sequence handles the full revenue workflow for complex pricing, from quoting and metering to invoicing, revenue recognition, and collections. Book a public demo at https://sequencehq.com and use code Cognizant in the source field to save 20% off year one AvePoint is building the control layer for AI agents so you can securely govern, audit, and recover every action at scale. Design trusted agentic outcomes from day one at https://avpt.co/tcr…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.