High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Kyle Corbitt: uncertainty

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“That's one thing. But I don't think any of that means that recursive self-improvement won't matter or doesn't matter.”

— Kyle Corbitt

Source trail

Everything needed to verify it.

Speaker
Kyle Corbitt
Attribution
Verified speaker
Claim type
uncertainty
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…But that would mean if a few things changed that You know, I guess obviously everybody's wondering, like, are we heading into recursive self-improvement? And if so, like, what's it going to mean? I've seen, you know, a bunch of papers probably from the kind of 18 to 36 months ago, vintage GPT-4 class models basically trying to do recursive self-improvement. And it seemed like, generally speaking, they would kind of get better for like three to five rounds and then kind of level off. And yet there's at least some expectation among people who've been right about a lot of things that this could go the other way in the not too distant future if models become smart enough to, I guess, maybe recursively self-improve in multiple ways, like not just critiquing their own outputs, but also just finding better architectures for themselves. And, you know, it could be a lot of different dimensions. in which they might self-improve. If that happens, you know, I mean, I also remember the Anthropic leaked pitch deck from a few years ago where they basically said, we think the people in 26 timeframe that train the best models might create such a big advantage that like nobody will ever catch up. Again, I've kind of filled in the gaps on that story for myself by thinking like, well, maybe it's these metacognitive behaviors, it's this sort of deeper understanding, problem solving ability, what have you. But you're kind of saying, nah, it's probably mostly compute and incentives and lack of inference business, which itself was very much related to compute. So I guess bottom line, it sounds like for you, if compute constraints were relaxed, you would expect to see Chinese companies be able to catch up and you wouldn't expect some sort of runaway dynamic to take hold where that would become impossible. Oh, I think that catching up right now is mostly compute gated. I mean, it's also like capital gated. I mean, in the sense that like buying the necessary compute certainly already requires billions of dollars and will require tens or hundreds of billions of dollars soon. So I think there's like an open question, like how healthy are the Chinese capital markets? Will they be able to make a case that they'll be able to keep their business if it goes really well, which I think has been a question with prior generations of Chinese tech companies, which might just be hard for them to overcome. That's one thing. But I don't think any of that means that recursive self-improvement won't matter or doesn't matter. My belief is that it probably does, and my belief is that we probably will reach it with the current generation or the next generation models. Because we already are in a self-improvement loop. That's what you have to remember, is these models keep getting better because we keep running more experiments and then figuring out, okay, what are the bottlenecks, let's solve those bottlenecks, and those happen at all levels, it happens at At the hardware level, figuring out like what's the most efficient way, algorithmic level, the data level, like these are all in self-improvement loops already. And there are multiple constraints, but one of the big constraints is just like human intelligence, right? Which is like, does... Are the people making those allocation decisions smart enough to make the right bets on what bottleneck to tackle next or what investments to make? And you can totally imagine that if you were just to staff OpenAI, if you just had a minimum bars, you're not allowed to be hired here unless you have an IQ of 180, I would imagine they would be able to solve those bottlenecks a lot faster if they could wave a mind wanting and get enough people that look like that. I don't know. I just feel like the bar for recursive self-improvement to take off is actually relatively low. I mean, it's just like, you just have to be better than the smartest human, which is not that smart. It's a wild time to be alive, that's for sure. And it does seem increasingly plausible that that could happen in the not too distant future. I don't know if you have anything more to say about recursive self-improvement. I was going to move next to the cottage industry of RL environment creation. I think this is kind of a, I don't know, people know it's out there, but it's kind of a dark matter sort of thing where because there's so few customers, it's not like these companies have much incentive to go talk super broadly about what they're doing. They probably, on the contrary, have the opposite, right? They know all the customers they can possibly sell to and telling the world more broadly what they're selling is just inviting competition that they don't want to have. It seems like the rest of us who aren't directly involved in the making, selling, and buying of these environments are kind of left in the dark. What can you tell me from what you've seen about that seemingly rapidly growing niche? Like, how big is it? Who's doing it? What do the environments look like? What makes a good environment? So on and so forth.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence