observation · 13 Feb 2026 · 6:59

Early AI models before GPT-1 were trained on datasets that lacked a wide distribution of text, relying on narrow benchmarks like fanfiction corpora.

You had very standard language modeling benchmarks. GPT-1 itself was trained on a bunch of fanfiction, I think actually. It was literary text, which is a very small fraction of the text you can get. In those days it was like a billion words or something, so small datasets representing a pretty narrow distribution of what you can see in the world.

Watch at 6:59