Evidence receipt / evaluation
Published · transcript-backedAndrej Karpathy: evaluation
17 Oct 2025 Dwarkesh Podcast Andrej Karpathy — AGI is still a decade away
“I guess they just don’t work as well empirically because right now the models are collapsed.”
Source trail
Everything needed to verify it.
- Speaker
- Andrej Karpathy
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 17 Oct 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…What is a solution to model collapse? There are very naive things you could attempt. The distribution over logits should be wider or something. There are many naive things you could try. What ends up being the problem with the naive approaches? That’s a great question. You can imagine having a regularization for entropy and things like that. I guess they just don’t work as well empirically because right now the models are collapsed. But I will say most of the tasks that we want from them don’t actually demand diversity. That’s probably the answer to what’s going on. The frontier labs are trying to make the models useful. I feel like the diversity of the outputs is not so much... Number one, it’s much harder to work with and evaluate and all this stuff, but maybe it’s not what’s capturing most of the value. In fact, it’s actively penalized. If you’re super creative in RL, it’s not good.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.