Evidence receipt / recommendation
Published · transcript-backedKyle Corbitt: recommendation
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“d yeah, we work with lots of customers that look like that and get them to smaller models that have good quality.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, okay, so I'll start by caveating that like, this is my day job, this is the business that I work in. And so, you know, I guess use that as a sort of, to appropriately calibrate, you know, how you take my, you know, my advice here. That said, like I do think it's very, I try to be well calibrated and not to let my biases, you know, influence the recommendations I give. So anyway, take that for what it's worth. Yeah, so in general, the way I answer that question, when someone comes to me and says, Hey, should I be using fine-tuning? And usually it's for RL, 'cause that's what we find. At least on the capabilities point of view, it's a strict superset in my experience of what you can get with SFT, although we also support SFT with our platform and with our team. But when someone comes to me and asks if they should do it, the first question is basically, what is the problem you're trying to solve, and how frustrated are you with the Frontier models? And if the situation you're in is actually the Frontier models like work pretty well and there's like maybe these small issues I want to solve with it, but like, yeah, it can get the job done, then my advice is you should just stick with that because there are real downsides if you're bringing model customization into your stack. The biggest downside is it is going to slow down your iteration loop. Like that is, you know, and we're working, like that's our biggest focus as a team actually is building tooling and automations to decrease increase that cost, but it is a real cost. It's going to take you extra time every time you want to change one of your models if you're customizing it. So you should only do it if you're running into a major pain. Now, what are the pains that we see most often where it actually does justify that cost? Today, the biggest one by a large margin is around latency. So we have a lot of customers that are in, you know, oftentimes it's customer support or, you know, inbound sales on the phone, voice dictation companies. So Willow and Whisper are both customers of ours. And generally the common thread there is if you try and use a Frontier model for one of these, you're just gonna have a bad, you'll give your customers a bad experience because it takes too long to respond. And so that forces you to move to a smaller model. I mean, there's the tricks you can do as well. But ultimately, there is a ceiling on how many tokens per second you can get out of an extremely large model. So you're forced to move to a smaller one. And then in many cases, when you do move to that smaller model, you find that the quality is not where you need it to be to give a good experience. So if you're in that situation, then it can make sense to bring in fine-tuning. And yeah, we work with lots of customers that look like that and get them to smaller models that have good quality. Now, once you've paid the cost of, hey, I am going to introduce this extra complexity. d yeah, we work with lots of customers that look like that and get them to smaller models that have good quality. Now, once you've paid the cost of, hey, I am going to introduce this extra complexity. What we find is like typically on customer metrics, like, you know, number of cases closed, things like that, you can, using reinforcement learning, get to a better place. So you can exceed the performance of the frontier models, which is really fun. And your costs are also typically much lower on a like per token basis. So those are the secondary advantages as well. But I would say what's driving the decision most often in the current environment is latency. Okay, cool. Great answer. I expected nothing less. What are the sort of range of tasks that people are coming to you for? You mentioned a couple, but on the homepage, I noticed that it says, use reinforcement learning to train reliable agents. And, you know, in those couple of examples, those weren't really agent examples, I'm wondering kind of what agents people are fine tuning models for today and like, and how kind of broad of a remit those agents have within the environments where they're put to work.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.