High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Gwern: evaluation

13 Nov 2024 Dwarkesh Podcast Gwern — Anonymous writer who predicted AI trajectory on $12K/year salary

“When you only can do a few instances, you would typically find that it doesn't work, and you would give up and you would go away and do something else.”

— Gwern

Source trail

Everything needed to verify it.

Speaker
Gwern
Attribution
Verified speaker
Claim type
evaluation
Recorded
13 Nov 2024
Publisher
Dwarkesh Podcast

Transcript context

…I remember in 2020, people were writing bestselling books about AI. It was definitely a thing people were talking about, but people were not noticing the most salient things in retrospect: LLMs, GPT-3, scaling laws. All these people who are talking about AI but missing this crucial crux, what were they getting wrong? I think for the most part they were suffering from two issues. First, they had not been paying attention to all of the scaling results before that which were relevant. They had not really appreciated the fact that, for example, AlphaZero was discovered in part by DeepMind doing Bayesian optimization on the hyperparameters and noticing that you could just get rid of more and more of the tree search as you went and you got better models. That was a critical insight, which could only have been gained by having so much compute power that you could afford to train many, many versions and see the difference that that made. Similarly, those people simply did not know about the Baidu paper on scaling laws in 2017, which showed that the scaling laws just keep going and going forever, practically. It should have been the most important paper of the year, but a lot of people just did not prioritize it. It didn't have any immediate implication, and so it sort of got forgotten. People were too busy discussing Transformers or AlphaZero or something to really notice it. So that was one issue. Another issue is that they shared the basic error I was making about algorithms being more important than compute. This was, in part, due to a systematic falsification of the actual origins of ideas in the research literature. Papers do not tell you where the ideas come from in a truthful manner. They just tell you a nice sounding story about how it was discovered. They don’t tell you how it’s actually discovered. So even if you appreciate the role of trial and error and compute power in your own experiment as a researcher, you probably just think, “Oh, I got lucky that way. My experience is unrepresentative. Over in the next lab, there they do things by the power of thought and deep insight.” Then it turns out that everywhere you go, compute and data, trial and error, and serendipity play enormous roles in how things actually happened. Once you understand that, then you understand why compute comes first. You can't do trial and error and serendipity without it. You can write down all these beautiful ideas, but you just can't test them out. Even a small difference in hyperparameters, or a small choice of architecture, can make a huge difference to the results. When you only can do a few instances, you would typically find that it doesn't work, and you would give up and you would go away and do something else. Whereas if you had more compute power, you could keep trying. Eventually, you hit something that works great. Once you have a working solution, you can simplify it and improve it and figure out why it worked and get a nice, robust solution that would work no matter what you did to it. But until then, you're stuck. You're just flailing around in this regime where nothing works. So you have this horrible experience going through the old deep learning literature and seeing all sorts of contemporary ideas people had back then, which were completely correct. where nothing works. So you have this horrible experience going through the old deep learning literature and seeing all sorts of contemporary ideas people had back then, which were completely correct. But they didn't have the compute to train what you know would have worked. It’s just tremendously tragic. You can look at things like ResNets being published back in 1988, instead of 2015. And it would have worked! It did work, but at such a small scale that it was irrelevant. You couldn't use it for anything real. It just got forgotten, so you had to wait until 2015 for ResNets to actually come along and be a revolution in deep learning. So that’s kind of the double bias of why you would believe that scaling was not going to work. You did not notice the results that were key, in retrospect, like the BigGAN scaling to 300 million images. There are still people today who would tell you with a straight face that GANs cannot scale past millions of images. They just don't know that BigGAN handled 300 million images without a sweat. If you don't know that, well you probably would easily think, “Oh, GANs are broken.” But if you do know that, then you think to yourself, “How can algorithms be so important when all these different generative architectures all work so well—as long as you have lots and lots of GPUs?” That's the common ingredient. You have to have lots and lots of GPUs.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence