High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Kyle Corbitt: evaluation

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“It's called the credit assignment problem, right? Because really what you want to do is you want to assign credit and wait just the key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution.”

— Kyle Corbitt

Source trail

Everything needed to verify it.

Speaker
Kyle Corbitt
Attribution
Verified speaker
Claim type
evaluation
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…bad action, you want to update your weights to make it do less, right? And also importantly, this is something that happens at an action-by-action basis. But your reward in PPO can be very long-term. So it could be at the end of a very long sequence of actions, you finally find out that commonly this was used with games. And so you would say, Hey, at the end of the game or after a minute of gameplay, what's my score? or something like that. So what PPO does is a few different things, and it's actually, of course, building on older work as well, there's an called Reinforce, which is trying to solve the same problem. PPO adds some extra terms to keep it stable, keep it in sort of like a trust region where you're kind of hopeful that the model hasn't changed too much as you're updating it. But the key The key thing that PPO does, and actually this is not unique to PPO, this is from older than PPO, but you want to calculate the advantage of every single action. So every single time it takes an action, you want to say, Hey, was this a good or bad action? And the way it does that is by actually training a couple of different models in parallel. So you have the policy model, which is just your normal model that's generating the actions. And then you have a separate model, which is called the value model or the critic model. And the value model is actually predicting, saying, Hey, based on the set of actions up to this point, what do I think the score is going to be in the long-term? Basically, it's predicting for this action, what do I believe is the value of this action? What impact will this action have on the score in the long-term? And it's predicting that for every single action in the sequence. And then eventually you do get to see like what the actual score is. And then basically, if the score ends up much higher than you expected, then you can say, oh, some of these actions clearly were much more valuable than we expected. So if it's like, hey, my critic model thought it would have a low score and it actually has a high score, then I want to like make it much more likely that I have a high or that this action happens in the future. Okay, now moving on to GRPO, the sort of like key difference here is instead of figuring out what the value of any specific action is, oh, actually, before I go into GRPO, I should mention this all translates directly into LLMs. And the translation that people do, people have tried actually a lot of different translation, but the one that most people do, and it's kind of like the simplest thing that works is every single token generated is an action, right? So we're using the exact same concepts as we were using before and just saying like, Hey, every token, the state up to that point is the full context and then this token is an action and the next token is another action. What we do with GRPO is it turns out that calculating, figuring out that value model and keeping it up to date is painful. ontext and then this token is an action and the next token is another action. What we do with GRPO is it turns out that calculating, figuring out that value model and keeping it up to date is painful. It's just tricky to get right. It's like another set of hyperparameters you have to tune on, Okay, we have to keep this model updated or else train doesn't work well. What GRPO did, and they actually were not the first ones to do this, but they get the critics, they're the first ones do it at scale and prove it worked well, is I said, Hey, we're just going to completely throw away the value model. And what we're going to do is we're going to try... So the way we're going to figure out whether a given action is basically like a trajectory of actions is better or worse than what the model would have done otherwise is we're just going to run a bunch of them in parallel. So we are going to, with the exact same setup, the same initial conditions, we're going to run whatever, four or eight or 512, there's lots of different, you know, like hyperparameters between here as well, different runs in parallel. And we're going to see how often the model succeeds and how often it fails. And the reason we want to do this is because Let's say we've thrown away the critic model and we do a single run through with GRPO and we get a score and the score is one. Hey, it got it right. You don't know from that run whether the model just would always get this right or this is one in a million times that it got it right. If you're just naively updating your model because it got it right, but it would've always got it right, it's a spurious correlation where it's just like, Hey, it made some random choices, choices didn't affect the score. at all because, you know, it just always would get it right. And if you're upweighting those random choices it made, then you're, you know, just like kind of moving around in a pretty random direction. So what GRPO lets you do is it says, okay, you know, the sort of like advantage that we allocate to each of these tokens is going to be based on how much better this run did than the average of really what you want to compare this to the average. If we ran the current model infinite times on this, how much better did this do than that average? Obviously, we're not going to run infinite times, so we approximate that by doing it n times. Then getting to the end of your question, which is you're right. When we're actually updating the model weights, we are doing this at a token-by-token basis. Somehow we have to say for every single token, No, we want to update the weight such that this token is more likely if the advantage is positive or less likely if the advantage is negative. And this is a big problem in reinforcement learning. It's called the credit assignment problem, right? Because really what you want to do is you want to assign credit and wait just the key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution. he key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution. And so the sort of like, I guess the key insight of GRPO is to sort of just like do a very unsatisfying thing and kind of just punt on that a little bit. It's not a full punt. So what you do is you look at how likely every token was to be produced, right? Because you're sampling at a high temperature when you're doing these. And so some tokens it's like that it produces are very common, some tokens are not very common. And basically what you do is you say, hey, if my got a high score, then I want to give more credit to the tokens that, just by random chance, were less common. Because I assume that the high score is much higher than the average score across the entire group, then I assume that it probably was because there was some rare thing that I did in this that I didn't do in other cases, and that rare thing led to me doing well. And the exact same thing, the obstruction, if I get much lower score than the average on the group, then the rare things are the things that I'm going to penalize the the most because I'm like, Hey, that's probably what put me there. Now, you could ask the question, There could be many rare tokens. If you've got tens of thousands of tokens that are reading trace, how do you decide which rare token is most important? And you don't. You just throw up your hands and you say, All the rare tokens get upvoted the same way. This is, like I said, a very unsatisfying answer. And I think that's one of the reasons why there was an almost 10-year gap between PPO that had this value model that tried to determine on a token basis and GRPO where it's like, Hey, we're just going to throw that all away because it feels wrong. It feels like it shouldn't work. In practice, it does, though.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence