Evidence receipt / evaluation
Published · transcript-backedEric Jang: evaluation
15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch
“For small data regimes, my experience is that ResNets still outperform transformers and give you more bang for the buck at lower budgets.”
Source trail
Everything needed to verify it.
- Speaker
- Eric Jang
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…While you’re erasing, another thing that was important for me to understand was about the MCTS data structure with nodes and children of nodes. This is done per move and reinstantiated once a move is made. A human makes a move, then the AI looks at this and runs a bunch of simulations to figure out what move to make next. A simulation is exploring one more node in the MCTS tree. Once you run 1,000 simulations, that informs the probability of what move to make next, as you’ll explain. That’s what you store. You choose the best move given those probabilities. You discard all of that, the next player makes a move, and you restart this process at the beginning of every move. Correct. One small addendum: you don’t discard all of that. You keep one thing behind that we’ll use later. Now that we have a basic intuition of how moves are made with search, we’re going to talk about how neural networks can speed this up by providing an analog to human intuition. There are two networks. There is the value network, which takes in a state and predicts, am I going to win or lose? It’s a binary classification problem. Then we have a policy network, which induces a distribution over good actions to take. I’m going to draw a one-dimensional flattened move distribution, but this is really a square grid. These are the probability distributions over good actions. Both of these are categorical classification problems. You can train this like any classifier with deep learning, cross-entropy loss, that kind of stuff. The specific architecture does not matter too much. I tried a few different architectures. Transformers work, ResNets work. For small data regimes, my experience is that ResNets still outperform transformers and give you more bang for the buck at lower budgets. But this may not always be true. Wait, why is that?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.