Speakers in the public record
Claim mix
evaluation 9belief 7prediction 3recommendation 2uncertainty 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
22 published records
“d really there was this other token that was much more important that did in fact need to change to get to the right answer, but that one's just mixed in with all these other random unrelated changes. That general intuition generalizes to other task shapes as well, including creative writing, where maybe there's two different ways to phrase this and they're both fine, and the model would have chosen one, and your creative team chose another, and they're both okay, and you don't really want to waste your model updates, because every time you update the weights, there's a potential for catastrophic forgetting, and sort of just off-target effects in general.”
- Publisher
- The Cognitive Revolution
“When you describe, you said more specifically, you know, something that not the model can't get right, but that it rarely gets right. That's key because when we do things like GRPO, the you've got to have at least one right answer, right, to be to have any sort of advantage.”
- Publisher
- The Cognitive Revolution
“Like one open question was sort of like, hey, how do we do length normalization? And this comes into if you have a trace that happens to be, so, so it's sort of the original math in GRPO actually structurally advantaged very long thinking traces and, and, and, you know, just like generations in general, just because, you know, it didn't normalize by the number of tokens.”
- Publisher
- The Cognitive Revolution
“It's called the credit assignment problem, right? Because really what you want to do is you want to assign credit and wait just the key tokens that like were critical to this going right and not upweight the tokens that like, you know, just like always would have been right and didn't really contribute anything to the solution.”
- Publisher
- The Cognitive Revolution
“I would say our typical complaint was probably most often that, and this would definitely vary through different generations, but more recently it was like able to do the job perfectly well, so to speak.”
- Publisher
- The Cognitive Revolution
“I expect we'll see much more locking down models to specific use cases, specific product surfaces for multiple reasons, but a big one being the, because it makes distillation harder, especially distillation in, out of domain areas that aren't within that product surface.”
- Publisher
- The Cognitive Revolution
“How do you relate this to what I think of as metacognitive behaviors. In the original R1 paper, there was this aha moment that they published.”
- Publisher
- The Cognitive Revolution
“If that happens, you know, I mean, I also remember the Anthropic leaked pitch deck from a few years ago where they basically said, we think the people in 26 timeframe that train the best models might create such a big advantage that like nobody will ever catch up.”
- Publisher
- The Cognitive Revolution
“When I think about like how much the weights change with fine tuning, I usually think of that as kind of more a function of like some sort of divergence penalty, some sort of tethering of the, you know, the model as it's evolving to the base, to the starting point.”
- Publisher
- The Cognitive Revolution
“I think that's basically the return of the PPO value model, right? That's I should think about that kind of the same way.”
- Publisher
- The Cognitive Revolution
“Where I think you can make a stronger case if you're at OpenAI or Anthropic or even some of these like robot labs that it's like, hey, we have that strong sense of mission and we're the most likely place to change the world and so the absolute best people will go there instead so I suspect that he will not have outsized success in these areas but anyway that's that's speculation as well.”
- Publisher
- The Cognitive Revolution
“I think that said, like, I think you probably can get to superhuman performance on a composite task, like, you know, a very complex math proof, even without surpassing, like, you know, reaching a place where it's like no human could possibly have understood this or generated this, right? It's like, I mean, I think, I think that like one thing the models are very, very good at is going out on these long expeditions and fishing trips, where it's going very, very deep down a specific rabbit hole, and maybe they'll take that rabbit hole further than any human would because we'll lose the...”
- Publisher
- The Cognitive Revolution
“At that point, I suspect for a creative writing task, you would end up in a position where you're better off using the frontier models.”
- Publisher
- The Cognitive Revolution
“I think the rollout to the physical world would be very slow, probably just because there's a lot of constraints there and the data efficiency is going to be low, so the ROI is going to be relatively low and likely Frontier Labs is going to be very concentrated on automating everything digital first and then eventually potentially, there's these long tail, these physical things which are annoying to work with.”
- Publisher
- The Cognitive Revolution
“Yeah, like, I don't know, like, at the same time, I'm kind of like on the record is like being very skeptical of the human data labeling business, which is sort of like the prior thing.”
- Publisher
- The Cognitive Revolution
“d yeah, we work with lots of customers that look like that and get them to smaller models that have good quality.”
- Publisher
- The Cognitive Revolution
“This is, like I said, a very unsatisfying answer. And I think that's one of the reasons why there was an almost 10-year gap between PPO that had this value model that tried to determine on a token basis and GRPO where it's like, Hey, we're just going to throw that all away because it feels wrong.”
- Publisher
- The Cognitive Revolution
“I'm not sure, that doesn't feel alien to me if I'm sort of introspecting my own chain of thought or, you know, just like a conversation with someone, like that behavior doesn't feel weird, it feels very natural, and obviously reinforcement learning is bringing it out, because it is also true that that's the kind of behavior that, in retrospect, it makes sense both that like, oh yeah, that makes sense, but it also makes sense like, oh, this would not naturally come up in the pre-training data all that often, because usually if you're writing something on the internet and you have a new idea, you're not gonna chain of thought put out, oh wait, I have this other idea, you're going to condense it and just put your final thinking there.”
- Publisher
- The Cognitive Revolution
“Okay, so first of all, like, yeah, I think the reason GRPO specifically, like that algorithm and that acronym, like, you know, very concretely took off was not necessarily because it was like a big quantum leap on what came before.”
- Publisher
- The Cognitive Revolution
“That's one thing. But I don't think any of that means that recursive self-improvement won't matter or doesn't matter.”
- Publisher
- The Cognitive Revolution
“The reason why this is sort of cottage industry shaped, I believe, is for a few reasons one is the labs actually do have at least a weak preference for having lots of different vendors because you want if one person creates five different environments they are likely going to make similar assumptions and similar shortcuts in how they do all of them and so the signal that the model will gain from mastering all those environments is more correlated than when you would like than you would like and the whole game here is you want the broadest diversity of environments so having different people working on it is better Another reason why it's sort of cottage industry shaped is because this is extremely hard to hire for.”
- Publisher
- The Cognitive Revolution
“My guess is probably not. So I think concretely for your task, if the trade-off you're making is, Hey, we're going to take an open-source model and use RL to try and make it better at this, versus, Hey, we're just going to take whatever the best off-the-shelf model is and do prompt engineering, and we're allowing ourselves to expand to the best frontier models.”
- Publisher
- The Cognitive Revolution