High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Andrej Karpathy: belief

17 Oct 2025 Dwarkesh Podcast Andrej Karpathy — AGI is still a decade away

“Then I think you get away with a much smaller model because it’s a much better dataset and you could train it on it.”

— Andrej Karpathy

Source trail

Everything needed to verify it.

Speaker
Andrej Karpathy
Attribution
Verified speaker
Claim type
belief
Recorded
17 Oct 2025
Publisher
Dwarkesh Podcast

Transcript context

…Yeah, but I’m surprised that in 10 years, given the pace… We have gpt-oss-20b. That’s way better than GPT-4 original, which was a trillion plus parameters. Given that trend, I’m surprised you think in 10 years the cognitive core is still a billion parameters. I’m surprised you’re not like, “Oh it’s gonna be like tens of millions or millions.” Here’s the issue, the training data is the internet, which is really terrible. There’s a huge amount of gains to be made because the internet is terrible. Even the internet, when you and I think of the internet, you’re thinking of like The Wall Street Journal. That’s not what this is. When you’re looking at a pre-training dataset in the frontier lab and you look at a random internet document, it’s total garbage. I don’t even know how this works at all. It’s some like stock tickers, symbols, it’s a huge amount of slop and garbage from like all the corners of the internet. It’s not like your Wall Street Journal article, that’s extremely rare. So because the internet is so terrible, we have to build really big models to compress all that. Most of that compression is memory work instead of cognitive work. But what we really want is the cognitive part, delete the memory. I guess what I’m saying is that we need intelligent models to help us refine even the pre-training set to just narrow it down to the cognitive components. Then I think you get away with a much smaller model because it’s a much better dataset and you could train it on it. But probably it’s not trained directly on it, it’s probably distilled from a much better model still. But why is the distilled version still a billion?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence