Evidence receipt / evaluation
Published · transcript-backedKyle Corbitt: evaluation
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“Okay, so first of all, like, yeah, I think the reason GRPO specifically, like that algorithm and that acronym, like, you know, very concretely took off was not necessarily because it was like a big quantum leap on what came before.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Sequence handles the full revenue workflow for complex pricing, from quoting and metering to invoicing, revenue recognition, and collections. Book a public demo at https://sequencehq.com and use code Cognizant in the source field to save 20% off year one AvePoint is building the control layer for AI agents so you can securely govern, audit, and recover every action at scale. Design trusted agentic outcomes from day one at https://avpt.co/tcr Let me take several parts of this question and I will finish on the one you were there at the end and then hopefully that'll give you the chance to ask follow-ups if my explanation doesn't make sense. Okay, so first of all, like, yeah, I think the reason GRPO specifically, like that algorithm and that acronym, like, you know, very concretely took off was not necessarily because it was like a big quantum leap on what came before. It was because DeepSeek did a lot of engineering work around actually scaling it and released an actual artifact, a model that worked really well with it. Like that was kind of the reason why, you know, there was a whole constellation of other algorithms that probably would have worked just about as well. There was one that came out a little bit before called RLOO, ARLU, which basically is the same as GRPO and likely would have worked just as well if you'd scaled it. After GRPO, very shortly after, within a few months, certainly of the R1 release, there were various numerous improvements made upon it, which really do probably deserve their own algorithms. So there was a paper called DAPO, there's GSPO came out from the Quinn Lab, I believe, and then CISPO was another one that came out shortly after that They're all significant improvements. And then there's a bunch of minor tweaks that don't even have named things. So I would say, yeah, the algorithm that people use today in practice is actually probably further away from GRPO as initially described as GRPO was from what came before it. But we all just still call it GRPO because that was kind of the name that stuck. The, okay, so moving on to kind of like how it actually works and like let's, yeah, I'll talk through, I think this will be helpful to build your intuition on, you know, how the advantages are calculated and everything. So maybe I'll talk first about what came before GRPO, because GRPO is kind of interesting in that like a big part of its like development was that it threw away something that everyone had used before and some people still use. which so it's sort of the spiritual grandfather of all RL that people do on LLMs is an algorithm called PPO that was developed by John Schulman in 2017 I believe actually pre LLMs or pre the being big it was used for games and stuff and the key thing about PPO is it's sort of like you have your policy which is what you call the model your training it's taking a bunch of actions and the key thing you need to do is every time it takes an action you have to like kind of score how good or this action is. If it's a good action, then you want to, you basically want to update your weights to make it more likely to do that action. And if it's a bad action, you want to update your weights to make it do less, right? And also importantly, this is something that happens at an action-by-action basis. But your reward in PPO can be very long-term. bad action, you want to update your weights to make it do less, right? And also importantly, this is something that happens at an action-by-action basis. But your reward in PPO can be very long-term. So it could be at the end of a very long sequence of actions, you finally find out that commonly this was used with games. And so you would say, Hey, at the end of the game or after a minute of gameplay, what's my score? or something like that. So what PPO does is a few different things, and it's actually, of course, building on older work as well, there's an called Reinforce, which is trying to solve the same problem. PPO adds some extra terms to keep it stable, keep it in sort of like a trust region where you're kind of hopeful that the model hasn't changed too much as you're updating it. But the key The key thing that PPO does, and actually this is not unique to PPO, this is from older than PPO, but you want to calculate the advantage of every single action. So every single time it takes an action, you want to say, Hey, was this a good or bad action? And the way it does that is by actually training a couple of different models in parallel. So you have the policy model, which is just your normal model that's generating the actions. And then you have a separate model, which is called the value model or the critic model. And the value model is actually predicting, saying, Hey, based on the set of actions up to this point, what do I think the score is going to be in the long-term? Basically, it's predicting for this action, what do I believe is the value of this action? What impact will this action have on the score in the long-term? And it's predicting that for every single action in the sequence. And then eventually you do get to see like what the actual score is. And then basically, if the score ends up much higher than you expected, then you can say, oh, some of these actions clearly were much more valuable than we expected. So if it's like, hey, my critic model thought it would have a low score and it actually has a high score, then I want to like make it much more likely that I have a high or that this action happens in the future. Okay, now moving on to GRPO, the sort of like key difference here is instead of figuring out what the value of any specific action is, oh, actually, before I go into GRPO, I should mention this all translates directly into LLMs. And the translation that people do, people have tried actually a lot of different translation, but the one that most people do, and it's kind of like the simplest thing that works is every single token generated is an action, right? So we're using the exact same concepts as we were using before and just saying like, Hey, every token, the state up to that point is the full context and then this token is an action and the next token is another action. What we do with GRPO is it turns out that calculating, figuring out that value model and keeping it up to date is painful.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.