High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Kyle Corbitt: evaluation

1 May 2026 The Cognitive Revolution The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

“I'm not sure, that doesn't feel alien to me if I'm sort of introspecting my own chain of thought or, you know, just like a conversation with someone, like that behavior doesn't feel weird, it feels very natural, and obviously reinforcement learning is bringing it out, because it is also true that that's the kind of behavior that, in retrospect, it makes sense both that like, oh yeah, that makes sense, but it also makes sense like, oh, this would not naturally come up in the pre-training data all that often, because usually if you're writing something on the internet and you have a new idea, you're not gonna chain of thought put out, oh wait, I have this other idea, you're going to condense it and just put your final thinking there.”

— Kyle Corbitt

Source trail

Everything needed to verify it.

Speaker
Kyle Corbitt
Attribution
Verified speaker
Claim type
evaluation
Recorded
1 May 2026
Publisher
The Cognitive Revolution

Transcript context

…So my other side by side, there is the aha moment where the model is solving some math problem and it realizes that the way it had been doing it, was flawed, but now it recognizes there's another way and it like kind of takes a step back and comes at it from another different direction. And clearly we are seeing in frontier models a lot more of this sort of persistent, resilient, try, try again, problem solving that again, it's like somewhere, you know, deep in the long tail of the internet, somebody's written out how to do that. So it's like a little bit in the pre-training. There's supervised fine tuning, at least sometimes in these recipes as well, where you could potentially try to seed the kind of metacognitive strategies that you want. And then it seems like reinforcement learning is doing a lot to bring that forward as well. How do you think about like really what's driving that? And are we seeing things that are kind of alien problem solving, should we expect to see, are we seeing, and should we expect to see sort of alien reasoning approaches that are kind of not inspired by humans emerging through RL over time? Yeah, I think that's an interesting question. You know, I, I personally don't really feel like the so-called aha moment or, you know, I think, you know, wait is one that shows up all the time, right? Where the models will say wait and that's sort of like a code to say, hey, let's explore another direction. I'm not sure, that doesn't feel alien to me if I'm sort of introspecting my own chain of thought or, you know, just like a conversation with someone, like that behavior doesn't feel weird, it feels very natural, and obviously reinforcement learning is bringing it out, because it is also true that that's the kind of behavior that, in retrospect, it makes sense both that like, oh yeah, that makes sense, but it also makes sense like, oh, this would not naturally come up in the pre-training data all that often, because usually if you're writing something on the internet and you have a new idea, you're not gonna chain of thought put out, oh wait, I have this other idea, you're going to condense it and just put your final thinking there. But I'm sure it comes up sometimes, where you're in a chat history or whatever, anyway. So I don't think that's surprising to me. I think there's a separate, so I would say short answer, I have not seen strong evidence yet where it's like, oh, they're thinking in ways that are totally foreign, totally alien, hard for us to introspect, or to follow as a human. Now, there's a separate question, like, will we see more of that? I think in the limit, it seems very likely that the sort of ideal form of cognition for these artifacts and and just, you know, the ideal form of cognition generally likely looks would be something that looks very alien to a human. And so as we put more effort into RL and perhaps come up with better techniques to explore more, you know, on that sort of like explore exploit spectrum, then it would not surprise me if we do start seeing more of that. But I haven't seen it yet. Yeah. I mean, it's a bit of a different dimension on which it might arise, but just in terms of an intuition of what that might look like, the coconut paper out of meta maybe a year ago or something where they, it was basically like thinking in latent space. So instead of cashing a forward pass out to a token, I forget exactly what the like decision mechanism was for when it would pass its last internal state back to the next position as an embedding versus when it would actually cash out a token. There was some decider mechanism there somewhere, but at least for a while it could and would just loop on its own internal states rather than emitting and appending a token. And they found that it was much better at like graph search type problems that benefited from the ability to parallelize. It seemed like it was able to effectively run multiple branches of, go down multiple paths in parallel, in latent space together, because it was able to like chew out these things rather than having to spit out one token. I get a little scared of those kinds of innovations, honestly, 'cause I kind of wanna know what my AIs are thinking and that doesn't really lend itself to that. The other one that comes to mind is like, and I've been quite confused about this too, you might be able to shed some light on it. Apollo Research, when they did, I think it was '03, maybe it was '01, testing got access to chain of thought and they reported that the chain of thought was starting to look kind of bizarre. You remember the like disclaim, disclaim, vantage, you know, that weird sort of internal... I kind of was thinking of it as a dialect and I had kind of assumed that there was maybe sort of a chain of thought length penalty. Like if the original GRPO was like accidentally rewarding long chains of thought, it would also stand to reason like computers scarce. We want to keep these chains of thought as tight as possible, but then maybe overdo that. And now you're just starting to see like weird dialects emerge. How much have I gone off the rails in telling myself that story?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence