High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Kyle Corbitt: prediction

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“I expect we'll see much more locking down models to specific use cases, specific product surfaces for multiple reasons, but a big one being the, because it makes distillation harder, especially distillation in, out of domain areas that aren't within that product surface.”

— Kyle Corbitt

Source trail

Everything needed to verify it.

Speaker
Kyle Corbitt
Attribution
Verified speaker
Claim type
prediction
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, that's a good question. And I guess, again, an empirical one. So a few different thoughts. One is the most natural way in my mind to use Frontier models to bootstrap your own, your own near frontier models with reinforcement, in general, is to use the frontier models as judges. They're very good at that and that sidesteps the issue that you can't actually get and train on the chain of thought traces directly. So if you just kind of have a standard, hey, we're going to use a frontier model as our rubric and we'll have our model do generations that get judged, that's a very productive way. And in the blog post that Anthropic made about the distillation attacks, as they called them, from Chinese models. They specifically called out, I mean, they didn't say the breakdown of what all these were being used for, but they did say that one of the uses that they included in their general bucket was using their model as an LMS judge for other outputs. So that's one way where, yes, I think very clearly you can use the existence of a high-quality frontier model to improve your own. And I think that the nice thing about that approach as well is both you get those benefits of like, hey, you're staying in your own distribution because you're just using it as a judge, you're not doing SFT. But also in general with RL, you can train the model undertraining to be better than the teacher model that way. So it is a path to getting frontier level or pushing the frontier, even if you aren't starting from a frontier model. This is true in our own experiments. This is clearly true from the Frontier Labs because we see OpenAI and others as well using their N minus one generation model as a judge when they're in the process of training the next version of models. So that's the most natural way. As far as like using distillation directly, like, you know, SFT style. Yeah, I'm sure that does happen. I would imagine that happens pretty at like a relatively low volume and fairly early in the process before you do RL and my guess is that it's like not that valuable and and you can really get it's it's a shortcut that lets you use less compute but like not orders of magnitude less compute relative to just doing RL. And so, and particularly like as we see frontier models start to shut down their APIs more, which I think is just like, you know, I think is the more interesting direction to investigate or to sort of explore. You know, like we're already seeing, of course, like starting with the reasoning models, we're not seeing all the tokens that are produced anymore. We, you know, there are certain models. Yeah, they're cutting, they're not letting you see all the log probs. They're certainly not letting you see the prompt log probs. You know, like certain models, like for weeks, OpenAI will only let you use their models through Codex or, you know, and I expect we'll see more of that over time, not less. see the prompt log probs. You know, like certain models, like for weeks, OpenAI will only let you use their models through Codex or, you know, and I expect we'll see more of that over time, not less. I expect we'll see much more locking down models to specific use cases, specific product surfaces for multiple reasons, but a big one being the, because it makes distillation harder, especially distillation in, out of domain areas that aren't within that product surface. So I guess translating that to expectations, one story you could tell, which I've kind of been telling myself recently, is like, why are the Chinese models spikier or more apparently benchmaxed or whatever? I had been kind of thinking, well, they're like probably doing a lot of supervised fine tuning on frontier model outputs and therefore they're maybe not developing some of these like more persistent problem solving metacognitive behaviors that really allow the model to generalize robustly out of domain, right? Like I might not care so much about that exact question. What I really care about is in the chain of thought, like how good is it at breaking down and, you know, coming at problems from lots of different directions, but you're, accounts so far has gone the other way. I'm not sure if it's the other way, but now I'm not quite sure. That story doesn't ring so true anymore if you're saying they're probably not doing that much supervised fine-tuning and it's relatively early in the process and it's a compute saver, sure, but it's not a huge difference maker. So what is the difference? Are they just not so good at RL or they just don't have so much compute? Why are the Chinese companies not able to match the American frontier companies right now?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence