High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Eric Jang: evaluation

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.”

— Eric Jang

Source trail

Everything needed to verify it.

Speaker
Eric Jang
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…But isn’t that how it’s written to begin with anyways? The sum of the rewards? The thing that’s a little bit hidden here in the math is that we’re assuming that when you decompose the problem into a multi-step problem, you’re now introducing correlations between your actions through the computation of this guy. If you separate these things out, this will magnify the variance of this one. In the case where you don’t separate it out, if you just have T=1, you have a single estimate of log prob and a single estimate of reward. This term still shows up. In LLMs, the naive REINFORCE estimator looks a bit like the return of the single action… It looks kind of like this. This is the basic form here, but it’s still a contributor to variance. You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better. Intuitively, the analogy is: can we find a term in our training objective such that it’s actually discouraged from doing this, or these don’t have any effect on the gradient, and this has an effect on the gradient? If you applied that there, the only thing you could do is eliminate 49 of the games. At least the way you have it written there, it would be 51 times……

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence