High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Kyle Corbitt

Published podcast speaker

Claims
16
Episodes
1
Shows
1
Named items
0

Claim ledger

What Kyle said.

8 transcript-backed records

01 / evaluation

d really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. That general intuition generalizes to other task shapes as well, including creative writing, where maybe there's two different ways to phrase this and they're both fine, and the model would have chosen one, and your creative team chose another, and they're both okay, and you don't really want to waste your model updates, because every time you update the weights, there's a potential for catastrophic forgetting, and sort of just off-target effects in general.

“d really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. That general intuition generalizes to other task shapes as well, including creative writing, where maybe there's two different ways to phrase this and they're both fine, and the model would have chosen one, and your creative team chose another, and they're both okay, and you don't really want to waste your model updates, because every time you update the weights, there's a potential for catastrophic forgetting, and sort of just off-target effects in general.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

02 / evaluation

Like one open question was sort of like, hey, how do we do length normalization? And this comes into if you have a trace that happens to be, so, so it's sort of the original math in GRPO actually structurally advantaged very long thinking traces and, and, and, you know, just like generations in general, just because, you know, it didn't normalize by the number of tokens.

“Like one open question was sort of like, hey, how do we do length normalization? And this comes into if you have a trace that happens to be, so, so it's sort of the original math in GRPO actually structurally advantaged very long thinking traces and, and, and, you know, just like generations in general, just because, you know, it didn't normalize by the number of tokens.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

03 / evaluation

It's called the credit assignment problem, right? Because really what you want to do is you want to assign credit and wait just the key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution.

“It's called the credit assignment problem, right? Because really what you want to do is you want to assign credit and wait just the key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

04 / evaluation

Yeah, like, I don't know, like, at the same time, I'm kind of like on the record is like being very skeptical of the human data labeling business, which is sort of like the prior thing.

“Yeah, like, I don't know, like, at the same time, I'm kind of like on the record is like being very skeptical of the human data labeling business, which is sort of like the prior thing.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

05 / evaluation

This is, like I said, a very unsatisfying answer. And I think that's one of the reasons why there was an almost 10-year gap between PPO that had this value model that tried to determine on a token basis and GRPO where it's like, Hey, we're just going to throw that all away because it feels wrong.

“This is, like I said, a very unsatisfying answer. And I think that's one of the reasons why there was an almost 10-year gap between PPO that had this value model that tried to determine on a token basis and GRPO where it's like, Hey, we're just going to throw that all away because it feels wrong.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

06 / evaluation

I'm not sure, that doesn't feel alien to me if I'm sort of introspecting my own chain of thought or, you know, just like a conversation with someone, like that behavior doesn't feel weird, it feels very natural, and obviously reinforcement learning is bringing it out, because it is also true that that's the kind of behavior that, in retrospect, it makes sense both that like, oh yeah, that makes sense, but it also makes sense like, oh, this would not naturally come up in the pre-training data all that often, because usually if you're writing something on the internet and you have a new idea, you're not gonna chain of thought put out, oh wait, I have this other idea, you're going to condense it and just put your final thinking there.

“I'm not sure, that doesn't feel alien to me if I'm sort of introspecting my own chain of thought or, you know, just like a conversation with someone, like that behavior doesn't feel weird, it feels very natural, and obviously reinforcement learning is bringing it out, because it is also true that that's the kind of behavior that, in retrospect, it makes sense both that like, oh yeah, that makes sense, but it also makes sense like, oh, this would not naturally come up in the pre-training data all that often, because usually if you're writing something on the internet and you have a new idea, you're not gonna chain of thought put out, oh wait, I have this other idea, you're going to condense it and just put your final thinking there.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

07 / evaluation

Okay, so first of all, like, yeah, I think the reason GRPO specifically, like that algorithm and that acronym, like, you know, very concretely took off was not necessarily because it was like a big quantum leap on what came before.

“Okay, so first of all, like, yeah, I think the reason GRPO specifically, like that algorithm and that acronym, like, you know, very concretely took off was not necessarily because it was like a big quantum leap on what came before.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution

08 / evaluation

The reason why this is sort of cottage industry shaped, I believe, is for a few reasons one is the labs actually do have at least a weak preference for having lots of different vendors because you want if one person creates five different environments they are likely going to make similar assumptions and similar shortcuts in how they do all of them and so the signal that the model will gain from mastering all those environments is more correlated than when you would like than you would like and the whole game here is you want the broadest diversity of environments so having different people working on it is better Another reason why it's sort of cottage industry shaped is because this is extremely hard to hire for.

“The reason why this is sort of cottage industry shaped, I believe, is for a few reasons one is the labs actually do have at least a weak preference for having lots of different vendors because you want if one person creates five different environments they are likely going to make similar assumptions and similar shortcuts in how they do all of them and so the signal that the model will gain from mastering all those environments is more correlated than when you would like than you would like and the whole game here is you want the broadest diversity of environments so having different people working on it is better Another reason why it's sort of cottage industry shaped is because this is extremely hard to hire for.”
Speaker
Kyle Corbitt
Publisher
The Cognitive Revolution
Search evidence