High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Kyle Corbitt: evaluation

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“This is, like I said, a very unsatisfying answer. And I think that's one of the reasons why there was an almost 10-year gap between PPO that had this value model that tried to determine on a token basis and GRPO where it's like, Hey, we're just going to throw that all away because it feels wrong.”

— Kyle Corbitt

Source trail

Everything needed to verify it.

Speaker
Kyle Corbitt
Attribution
Verified speaker
Claim type
evaluation
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…ontext and then this token is an action and the next token is another action. What we do with GRPO is it turns out that calculating, figuring out that value model and keeping it up to date is painful. It's just tricky to get right. It's like another set of hyperparameters you have to tune on, Okay, we have to keep this model updated or else train doesn't work well. What GRPO did, and they actually were not the first ones to do this, but they get the critics, they're the first ones do it at scale and prove it worked well, is I said, Hey, we're just going to completely throw away the value model. And what we're going to do is we're going to try... So the way we're going to figure out whether a given action is basically like a trajectory of actions is better or worse than what the model would have done otherwise is we're just going to run a bunch of them in parallel. So we are going to, with the exact same setup, the same initial conditions, we're going to run whatever, four or eight or 512, there's lots of different, you know, like hyperparameters between here as well, different runs in parallel. And we're going to see how often the model succeeds and how often it fails. And the reason we want to do this is because Let's say we've thrown away the critic model and we do a single run through with GRPO and we get a score and the score is one. Hey, it got it right. You don't know from that run whether the model just would always get this right or this is one in a million times that it got it right. If you're just naively updating your model because it got it right, but it would've always got it right, it's a spurious correlation where it's just like, Hey, it made some random choices, choices didn't affect the score. at all because, you know, it just always would get it right. And if you're upweighting those random choices it made, then you're, you know, just like kind of moving around in a pretty random direction. So what GRPO lets you do is it says, okay, you know, the sort of like advantage that we allocate to each of these tokens is going to be based on how much better this run did than the average of really what you want to compare this to the average. If we ran the current model infinite times on this, how much better did this do than that average? Obviously, we're not going to run infinite times, so we approximate that by doing it n times. Then getting to the end of your question, which is you're right. When we're actually updating the model weights, we are doing this at a token-by-token basis. Somehow we have to say for every single token, No, we want to update the weight such that this token is more likely if the advantage is positive or less likely if the advantage is negative. And this is a big problem in reinforcement learning. It's called the credit assignment problem, right? Because really what you want to do is you want to assign credit and wait just the key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution. he key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution. And so the sort of like, I guess the key insight of GRPO is to sort of just like do a very unsatisfying thing and kind of just punt on that a little bit. It's not a full punt. So what you do is you look at how likely every token was to be produced, right? Because you're sampling at a high temperature when you're doing these. And so some tokens it's like that it produces are very common, some tokens are not very common. And basically what you do is you say, hey, if my got a high score, then I want to give more credit to the tokens that, just by random chance, were less common. Because I assume that the high score is much higher than the average score across the entire group, then I assume that it probably was because there was some rare thing that I did in this that I didn't do in other cases, and that rare thing led to me doing well. And the exact same thing, the obstruction, if I get much lower score than the average on the group, then the rare things are the things that I'm going to penalize the the most because I'm like, Hey, that's probably what put me there. Now, you could ask the question, There could be many rare tokens. If you've got tens of thousands of tokens that are reading trace, how do you decide which rare token is most important? And you don't. You just throw up your hands and you say, All the rare tokens get upvoted the same way. This is, like I said, a very unsatisfying answer. And I think that's one of the reasons why there was an almost 10-year gap between PPO that had this value model that tried to determine on a token basis and GRPO where it's like, Hey, we're just going to throw that all away because it feels wrong. It feels like it shouldn't work. In practice, it does, though. And is the intuition there kind of like, I studied this a little bit, but not in enough depth to be confident, but I'm sort of imagining that as we go through a chain of thought, there are critical tokens where you're either taking the right path or the wrong path, and then there are probably a bunch of tokens that kind of follow once you've made that critical decision that are all kind of naturally going to follow because that's just the structure of language. So you're trying to isolate in saying the ones that the model was least confident about, you're trying to zoom in or isolate or focus, if not isolate, emphasize maybe the right word, the critical decision points in that trace.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence