High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Gwern: evaluation

13 Nov 2024 Dwarkesh Podcast Gwern — Anonymous writer who predicted AI trajectory on $12K/year salary

“I thought it was ludicrous to suggest that simply because there’s some supercomputer out there which matches the human brain, then that would just summon out of nonexistence the correct algorithm.”

— Gwern

Source trail

Everything needed to verify it.

Speaker
Gwern
Attribution
Verified speaker
Claim type
evaluation
Recorded
13 Nov 2024
Publisher
Dwarkesh Podcast

Transcript context

…You're one of the only people outside OpenAI in 2020 who had a picture of the way in which AI was progressing and had a very detailed theory, an empirical theory of scaling in particular. I’m curious what processes you were using at the time which allowed you to see the picture you painted in the “Scaling Hypothesis” post that you wrote at the time. If I had to give an intellectual history of that for me, it would start in the mid-2000s when I’m reading Moravec and Ray Kurzweil. At the time, they're making this kind of fundamental connectionist argument that if you had enough computing power, that could result in discovering the neural network architecture that matches the human brain. And that until that happens, until that amount of computing power is available, AI is basically futile. To me, I found this argument very unlikely, because it’s very much a “build it and they will come” view of progress, which at the time I just did not think was correct. I thought it was ludicrous to suggest that simply because there’s some supercomputer out there which matches the human brain, then that would just summon out of nonexistence the correct algorithm. Algorithms are really complex and hard! They require deep insight—or at least I thought they did. It seemed like really difficult mathematics. You can't just buy a bunch of computers and expect to get this advanced AI out of it! It just seemed like magical thinking. So I knew the argument, but I was super skeptical. I didn't pay too much attention, but Shane Legg and some others were very big on this in the years following. And as part of my interest in transhumanism and LessWrong and AI risk, I was paying close attention to Legg’s blog posts where he's extrapolating out the trend with updated numbers from Kurzweil and Moravec. And he's giving very precise predictions about how we’re going to get the first generalist system around 2019, as Moore's law keeps going. And then around 2025, we'll get the first human-ish agents with generalist capabilities. Then by 2030, we should have AGI. Along the way, DanNet and AlexNet came out. When those came out I was like, “Wow, that's a very impressive success story of connectionism. But is it just an isolated success story? Or is this what Kurzweil and Moravec and Legg were predicting— that we would get GPUs and then better algorithms would just show up?” So I started thinking to myself that this is something to keep an eye on. Maybe this is not quite as stupid an idea as I had originally thought. I just keep reading deep learning literature and noticing again and again that the dataset size keeps getting bigger. The models keep getting bigger. The GPUs slowly crept up from one GPU—the cheapest consumer GPU—to two, and then they were eventually training on eight. And you can just see the fact that the neural networks keep expanding from these incredibly niche use cases that do next to nothing. The use just kept getting broader and broader and broader. I would say to myself, “Wow, is there anything CNNs can't do?” I would just see people apply CNN to something else every individual day on arXiv. So for me it was this gradual trickle of drops hitting me in the background as I was going along with my life. Every few days, another drop would fall. I’d go, “Huh? lse every individual day on arXiv. So for me it was this gradual trickle of drops hitting me in the background as I was going along with my life. Every few days, another drop would fall. I’d go, “Huh? Maybe intelligence really is just a lot of compute applied to a lot of data, applied to a lot of parameters. Maybe Moravec and Legg and Kurzweil were right.” I’d just note that, and continue on, thinking to myself, “Huh, if that was true, it would have a lot of implications.” So there was no real eureka moment there. It was just continually watching this trend that no one else seemed to see, except possibly a handful of people like Ilya Sutskever, or Schmidhuber. I would just pay attention and notice that the world over time looked more like their world than it looked like my world, where algorithms are super important and you need like deep insight to do stuff. Their world just kept happening. And then GPT-1 comes out and I was like, “Wow, this unsupervised sentiment neuron is just learning on its own. That's pretty amazing.” It was also a very compute-centric view. You just build the Transformer and the intelligence will come. And then GPT-2 comes out and I had this “holy shit!” moment. You look at the prompting and the summarization: “Holy shit, do we live in their world? And then GPT-3 comes out and that was the crucial test. It's a big, big scale-up. It's one of the biggest scale-ups in all neural network history. Going from GPT-2 to GPT-3, that's not a super narrow specific task like Go. It really seemed like it was the crucial test. If scaling was bogus, then the GPT-3 paper should just be unimpressive and wouldn't show anything important. Whereas if scaling was true, you would just automatically be guaranteed to get so much more impressive results out of it than GPT-2. I opened up the first page, maybe the second page, and I saw the few-shot learning chart. And I'm like, “Holy shit, we are living in the scaling world. Legg and Moravec and Kurzweil were right!” And then I turned to Twitter and everyone else was like, “Oh, you know, this shows that scaling works so badly. Why, it's not even state-of-the-art!” That made me so angry I had to write all this up. Someone was wrong on the Internet.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence