High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Dwarkesh Podcast / episode intelligence

Eric Jang – Building AlphaGo from scratch

15 May 2026 27 published claims 2 attributable people

Speakers in the public record

Claim mix

evaluation 9belief 9prediction 3recommendation 2commitment 2observation 1uncertainty 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

27 published records

02 / evaluation

There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.

“There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

11 / belief

I think that’s where a lot of the lateral thinking ability of humans comes from: knowing when the track you’re pursuing or the objective you’re pursuing is not right, and you should be asking a different question.

“I think that’s where a lot of the lateral thinking ability of humans comes from: knowing when the track you’re pursuing or the objective you’re pursuing is not right, and you should be asking a different question.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

14 / belief

Certainly, LLMs manage to do something that looks like real human reasoning without having to do an explicit tree structure. That being said, I think the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback, even if not in exactly the same instantiation as AlphaGo.

“Certainly, LLMs manage to do something that looks like real human reasoning without having to do an explicit tree structure. That being said, I think the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback, even if not in exactly the same instantiation as AlphaGo.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

16 / evaluation

You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.

“You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

19 / prediction

Because ln(n) grows slower than n, over time you will move from the argmax being dominated by the exploration term, which is the second term here, to the argmax being dominated by the Q term, which is when you’ve done enough simulations and are confident that this is the branch to go down.

“Because ln(n) grows slower than n, over time you will move from the argmax being dominated by the exploration term, which is the second term here, to the argmax being dominated by the Q term, which is when you’ve done enough simulations and are confident that this is the branch to go down.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

20 / commitment

For a long time, we made the assumption that if we can’t model the dynamics with a world model, we’re going to instead just collect trajectories and then plan with respect to the only number that really matters, which is reward.

“For a long time, we made the assumption that if we can’t model the dynamics with a world model, we’re going to instead just collect trajectories and then plan with respect to the only number that really matters, which is reward.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

22 / evaluation

In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.

“In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

24 / evaluation

So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.

“So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

26 / recommendation

If you just take this policy recommendation and take the argmax over these probabilities—if you take the argmax and just take this action as your Go play—it’ll be a very fast Go player that doesn’t think in terms of reasoning steps.

“If you just take this policy recommendation and take the argmax over these probabilities—if you take the argmax and just take this action as your Go play—it’ll be a very fast Go player that doesn’t think in terms of reasoning steps.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast
Search evidence