Evidence receipt / evaluation
Published · transcript-backedKyle Corbitt: evaluation
1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
“The reason why this is sort of cottage industry shaped, I believe, is for a few reasons one is the labs actually do have at least a weak preference for having lots of different vendors because you want if one person creates five different environments they are likely going to make similar assumptions and similar shortcuts in how they do all of them and so the signal that the model will gain from mastering all those environments is more correlated than when you would like than you would like and the whole game here is you want the broadest diversity of environments so having different people working on it is better Another reason why it's sort of cottage industry shaped is because this is extremely hard to hire for.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Corbitt
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, no, I can definitely speak to that. Yes, I have several friends who are founders of companies doing that, which is not saying much 'cause it feels like half the companies starting the last six months are doing that. So yeah, I mean, I think it's an interesting industry. Yeah, the general shape is you come up with some task that seems like it might be economically valuable, and usually it's these companies proposing the tasks to the labs, it's usually not the labs kind of like coming out and saying, Hey, we want, We want a shape like this, although that can happen as well. And so you try and come up with some task. And the trick is you want to package it up as something that is sort of agent-shaped. All of the dependencies can all be enclosed. You want to make sure that it's something that is very either ideally snapshot-able. So that's sort of the gold standard is something where it's like, at any point, you can kind of snapshot it and you can continue from that point. And then, of course, something that can be easily graded. And I've seen these where sometimes you have your own, like obviously the ideal thing is if you have sort of a gold standard of what the grade should be, a lot of these do end up are just not things that you can score in some absolute way. And so in those cases, usually the company will say, Hey, this is sort of like the rubric you have to grade. I've also heard that sometimes labs will just ignore those rubrics and they'll do their own rubrics internally because they think they they have better information on what good looks like. And so, yeah, these are things like, I mean, like lots of different web flows. So computer use, browser use, building, you know, copies, of course, of all the big apps. So you're getting copies of Jira and GitHub and, you know, flight booking and, you know, like office suites, like Google Sheets. You're trying to build environments that copy these. And then you're building that environment, and so that's all the dependencies, like the database, which is usually it's like SQLite or something, you want something ephemeral, and then you're also building the scores. And then the way it's deployed varies a lot as well. Even within a specific company, sometimes it can vary or with a specific lab. So sometimes the labs are like, require you to ship it all up in kind of a container they can run on their infrastructure. Other labs are fine with you running it yourself and they will just call your environment and, you know, just like, you know, run it and then you just give them the scores back. their infrastructure. Other labs are fine with you running it yourself and they will just call your environment and, you know, just like, you know, run it and then you just give them the scores back. The reason why this is sort of cottage industry shaped, I believe, is for a few reasons one is the labs actually do have at least a weak preference for having lots of different vendors because you want if one person creates five different environments they are likely going to make similar assumptions and similar shortcuts in how they do all of them and so the signal that the model will gain from mastering all those environments is more correlated than when you would like than you would like and the whole game here is you want the broadest diversity of environments so having different people working on it is better Another reason why it's sort of cottage industry shaped is because this is extremely hard to hire for. You know, this is, it is, it's sort of like a piecework style task where, you know, you're sort of like doing, building one environment, then you're building another, but the skill bar to doing this successfully is quite high, you have to put yourself in the... This is something we do internally for our customers all the time as we're building these environments at Coreweave, which we then use to train models. So I have trouble hiring people who can do a good job on this. Candidly, it's a very upper percentile engineer who's able to think through this in a way that actually gets it. And you don't even know if you got it wrong until way later in the process when you've trained a model with it and it's like, oh, did the model learn life skills or learn some hack on how to just get a high score. So there's a lot to keep in your head as you're doing this. And the people who are good at that are like, by definition, I mean, they're by definition smart and frontier adjacent, and they might just start a competitor to do this instead of joining you as an employee. So it becomes very, very difficult to scale. And also, the environments themselves are not a super durable resource in the sense that all these things get saturated fairly quickly. And so you really have to keep, you can't just like keep reselling the same environment to the same lab. Like they're probably going to be like, hey, that environment, you know, for the next model is already like the model can ace it and you have to just keep creating new ones. That's, yeah, it's fascinating. This may be hard to summarize and I don't know if anybody, you know, has enough outside of the labs, I guess would have enough information to really characterize this, but like, Is this a good business to be in? I can see it kind of going either way. I would assume if you've got a good environment, all the labs want to buy it. But then at the same time, they're buying a ton of stuff. How much does your one random thing add to the whole mess of things they already have? And also, it's depreciating, as you said, for you, right? So you've got to strike a deal. before they already saturate your thing and then truly don't need it anymore. Would you say this is a hot, good place for up and comers to go, or would you steer people away from it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.