High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“I think that's basically the return of the PPO value model, right? That's I should think about that kind of the same way.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, I mean, well, yes, it seems like it will clearly be necessary. There needs to be at some point you have to close the loop and get feedback from the real world. That process is much slower just naturally than anything digital, which is why we've seen, the real reason we've seen way more progress on the digital side is just because like those, you know, it's just much easier to gather the data, much easier to build the environments, everything is simpler, I think. But yeah, as we move past the digital realm into more physical things, yeah, clearly there will need to be data and training on that. What's, like, it's not totally clear to me what the shape will look like, and that'll be an interesting. You could imagine kind of like fully in the loop reinforcement learning where it's like, hey, we're trying some chemical reaction and then reading the data from it and then trying a new one and reinforcing on that directly. You could also imagine much more investment in AlphaFold style things where it's like, hey, we're just using the data to build really high quality simulations or world models of this specific area and then using those for RL and I kind of suspect that's where more of it will go. But even in that case, you still a lot of, you know, of the real world data to ground that simulation in. I think it'll be a very big business. I think that's basically the return of the PPO value model, right? That's I should think about that kind of the same way. Yes. Yeah, yeah, yeah, yeah. I mean, yeah, you can squint. You can definitely squint and say like, yeah, like a world model and a value model can, you know, serve, serve similar purposes.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence