High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Eric Jang: evaluation

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.”

— Eric Jang

Source trail

Everything needed to verify it.

Speaker
Eric Jang
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…Pedantic question, but is there a guarantee that MCTS must be better than the policy? For example, you could imagine early on in training, because MCTS is informed by the value network, that when the value network hasn’t been well-trained on finished games, MCTS is worse than a randomly initialized policy. Is it just a heuristic that MCTS is better than the policy, or is there some guarantee? In practice, it is a heuristic. It does work in practice, but let me illustrate an example where MCTS can give you a worse distribution than your policy network. This can often happen if your self-play algorithm has trained to a good point, but then somehow it collapses because it’s not trained on diverse data or something. Let’s say we have a board state where the policy recommendations are very good. So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans. Like in the case we showed with the corner play, maybe 100% of our training data in our replay buffer has lost examples of how to evaluate the value function at those states. You might end up in a scenario where your terminal value is very bad. If the terminal values of the leaves are not good, then this will propagate all the way up and cause your PUCT selection criteria and your backups to be off. And then you end up visiting a very different distribution than what your policy initially recommended. Also, if your number of sims is low, you might have a variance issue where you just don’t explore enough. It’s only guaranteed to converge when you take N to infinity. Variance in your search process as well as inaccuracies in your evaluation can definitely screw with the quality of your policy recommendation. That’s why it’s not a guaranteed improvement. That is why I suspect AlphaGo Lee had the playouts to the end in their training algorithm, so that they could ground this thing in real playouts. In practice, what you could also do is, for 10% of the games, prevent the bots from resigning and just say, “Resolve it to the end.” That way you get some training data in your replay buffer to really resolve those late-stage playouts that normal human players would not play to. If you assume that the value functions are correct, this is why MCTS gives you a better policy. It’s a very critical chain of assumptions. Assuming this is accurate, your search process should give you a better recommendation than your initial guess. If you have a cold-started policy—if you have an AlphaZero-type thing—really what’s happening for the first few epochs is that the policy is kind of useless. What you’re really just doing is, “Hey, let’s play full games, and once we have played full games, for the preceding moves, we’ll have labeled who won and who didn’t win.” The loss for AlphaZero has two components: how good is the policy relative to MCTS, and how good is the value prediction relative to who actually won the game from this move. You can think of this being applied to every single action or every single move. Really what’s happening at the beginning of AlphaZero training is just that we’re trying to get the value function to actually predict who will win the game if you find yourself in this state and you’re this player. Functionally, that’s all that’s happening. Later on, once that’s well trained, now the policy is also improving.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence