Evidence receipt / prediction
Published · transcript-backedEric Jang: prediction
15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch
“I was thinking, “Let’s see if the Bitter Lesson had happened, where a lot of these tricks just go away because Nvidia made faster GPUs.”
Source trail
Everything needed to verify it.
- Speaker
- Eric Jang
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 15 May 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…How much of the improvements to compute efficiency are methods that did not exist as of 2017 versus things which they could have done in 2017 but didn’t? Great question. Going into this project, I knew in the back of my mind that things always get easier to do over time. I wanted to see where Go was at, given that there hadn’t seemed to be any major open-source strong bot after KataGo in 2020. Reading the KataGo paper, there were a lot of clever ideas. I was thinking, “Let’s see if the Bitter Lesson had happened, where a lot of these tricks just go away because Nvidia made faster GPUs. Roughly, where are we on that?” Again, this is not a peer-reviewed claim. It’s just my preliminary vibe guess based on what I’ve seen with my own experiments. It seems like architecture choices don’t matter that much. Transformer versus ResNet… We’re at the speed of GPU where the size of the model is not so big that this really matters. You can simplify the setup quite a lot. Instead of doing a distributed asynchronous RL setup with replay buffers and pushers and collectors, you can do a dumb synchronous thing where you collect, train a supervised learning model, and then collect again. There are opportunities to simplify infrastructure. Nvidia GPUs have indeed gotten faster. Whereas KataGo was trained on V100s, you can train on half the number of desktop Blackwell GPUs and it still works. Some of the auxiliary supervision objectives that KataGo developed aren’t really necessary if you have a strong initialization. If you’re initializing best-response training against KataGo itself, your own model needs none of the tricks that KataGo needs. The core thing is getting as quickly as possible to some strong opponents. That matters a lot more than the specific architectural innovations. But there are still some nice compute multipliers. I found that training on 9x9 boards was very nice for resolving endgame value functions. If you can co-train that on an architecture that can transfer between 9x9 and 19x19, you can really cut down the warm-start time to learn from scratch. AlphaGo Zero’s plot showed the first 30 hours or so spent basically catching up to the supervised learning baseline. You can cut down that time a lot by pre-training on a small board and then warm-starting that into your 19x19 board play. There was some other stuff, like varying the number of sims between episodes. This turns out to be not that sensitive. You can fix it or increase it. It doesn’t matter too much. Anyway, it’s nice from a scientific perspective, revisiting an old paper and seeing what really matters. This is a tangential question, but why is it okay to have a buffer in AlphaGo? Every time I talk to an AI researcher, they’re telling me about how bad it is to be off-policy. But the way a naive implementation of AlphaGo Zero would work is that most of the moves in a given backward step, or in a batch of backward steps, would not be among the ones made by the most recently trained model. So why is that okay?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.