Evidence receipt / evaluation
Published · transcript-backedEric Jang: evaluation
15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch
“You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.”
Source trail
Everything needed to verify it.
- Speaker
- Eric Jang
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…But isn’t that how it’s written to begin with anyways? The sum of the rewards? The thing that’s a little bit hidden here in the math is that we’re assuming that when you decompose the problem into a multi-step problem, you’re now introducing correlations between your actions through the computation of this guy. If you separate these things out, this will magnify the variance of this one. In the case where you don’t separate it out, if you just have T=1, you have a single estimate of log prob and a single estimate of reward. This term still shows up. In LLMs, the naive REINFORCE estimator looks a bit like the return of the single action… It looks kind of like this. This is the basic form here, but it’s still a contributor to variance. You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better. Intuitively, the analogy is: can we find a term in our training objective such that it’s actually discouraged from doing this, or these don’t have any effect on the gradient, and this has an effect on the gradient? If you applied that there, the only thing you could do is eliminate 49 of the games. At least the way you have it written there, it would be 51 times……
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.