Evidence receipt / evaluation
Published · transcript-backedEric Jang: evaluation
15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch
“There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.”
Source trail
Everything needed to verify it.
- Speaker
- Eric Jang
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…The other question is how stackable local improvements are in the attempt to get to a better result on the outer loop. I’ve heard rumors that at some AI labs, the thing that has gone wrong is that people will individually pursue good ideas, but those don’t end up stacking well, and so the training run fails because of some weird interaction between two seemingly good ideas. Having a single top-down vision of how things should work is very important. Having worked at different AI labs and also played around with parallel agents trying different ideas, what’s your sense of how parallelizable AI innovation is? Great question. I think the research taste for executing well on the Bitter Lesson is that you need to know how much the Bitter Lesson can buy you and how much is too much to ask for, at any given moment. Of course, in the fullness of time compute is the single most important determinant of how things work. It’s almost inevitable that as you scale up energy and compute and parameters, intelligence will just fall out of that. That’s super beautiful, super profound. No algorithmic detail really matters beyond that. But in the present day, we don’t have infinite compute and parameters and arbitrarily good initialization, so we have to come up with heuristics that give us that. These heuristics are probably somewhat redundant. That’s probably why you see this effect where a lot of these compute multipliers don’t necessarily stack. They might have some correlated benefit. And then three years down the line, when the Nvidia GPUs have gotten even stronger, maybe they stack even less well. Maybe at any given point in time, the benefit of any given compute multiplier is transitory, which is what I suspected with the KataGo paper. There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else. I think that’s a matter of taste in the present time. How about the outer loop? How verifiable for making AI smarter? With Go, you do have this outer loop of win rate against the best open-source model out there. Even there, as you were saying, there are other outer loops of whether you discovered a new phenomenon, which is actually very hard to…. If you didn’t know scaling laws were important… When were Chinchilla or Kaplan scaling laws released?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.