Evidence receipt / prediction
Published · transcript-backedKyle Corbitt: prediction
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“My guess is probably not. So I think concretely for your task, if the trade-off you're making is, Hey, we're going to take an open-source model and use RL to try and make it better at this, versus, Hey, we're just going to take whatever the best off-the-shelf model is and do prompt engineering, and we're allowing ourselves to expand to the best frontier models.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Rarely at best, I would say, yeah. Yeah. So here's what I think. I think it is likely that you would have been able to get better performance out of the models with reinforcement learning than with SFT. And there's a few different factors here that sort of muddle it. I mean, one is that OpenAI's support for RL was half-hearted at best at any given point. And I think technically they still do it, but that entire model customization platform feels very much in maintenance mode at this point generally. So I think on that side, yeah, That might just not have worked. In a parallel universe where you were using an open-source model and you're using a QEN model or something like that, then I would say with a fairly high degree of confidence that if you're able to get decent results out of SFT, the ceiling of the best results you can get with reinforcement learning is going to be higher. And that's true even if the data you're using for SFT is high-quality data, human data. And the reason why is because it's just the whole trick to RL, the whole reason RL works, or is something that people invest in at all is because it turns out it really does matter how well your data distribution matches kind of the models, you know, like a standard mode of thinking or just like what it's picked up from pre-training. And what RL gets you is it just gets you, you know, more, it is working within those channels that are already carved quite deeply within the model. And when you work within those channels, you can just get a lot further because you're not trying to over overwrite what it's doing. And yeah, you might say, Well, overwriting is what we're trying to do. We're trying to get it to do something it's not good at, which is fair, but it ends up being quite destructive. It's actually really interesting if you look at the weights. If you're doing SFT, it's just like, Even with very few examples and even with a very low learning rate, it's just throwing the weights all to pieces and the average differences are so much larger than doing RL. That's a big part of why you get this catastrophic forgetting because it's overriding other pathways. It's just like you're trying to get the model to do something that's quite different than what it was trained to do, whereas RL is going to let you stay in those grooves and get a lot further. Yes, I do think that would have worked. Now, in your specific case, would it be worth it? Would it get you to a place where it's like, oh, this is better than just using the frontier? My guess is probably not. So I think concretely for your task, if the trade-off you're making is, Hey, we're going to take an open-source model and use RL to try and make it better at this, versus, Hey, we're just going to take whatever the best off-the-shelf model is and do prompt engineering, and we're allowing ourselves to expand to the best frontier models. At that point, I suspect for a creative writing task, you would end up in a position where you're better off using the frontier models. we're allowing ourselves to expand to the best frontier models. At that point, I suspect for a creative writing task, you would end up in a position where you're better off using the frontier models. And yeah, we can sort of get into-- there are definitely tasks where I would say the exact opposite and say that the RL could do well. I would also say that this is obviously all dependent on the amount of compute. I think, theoretically, anything's possible. If you buy yourself a data center and spend a couple billion dollars on this task, you would be able to surpass the frontier. But the trade-off point would be fairly long, I suspect, along that curve for a task of this shape.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.