Evidence receipt / recommendation
Published · transcript-backedEric Jang: recommendation
15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
Source trail
Everything needed to verify it.
- Speaker
- Eric Jang
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 15 May 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…How would you design a baseline where it’s truly better? This is where in RL people use things like TD learning to better approximate the quality function, the Q that we mentioned earlier. You can try to subtract that from your return. Ideally, what you really want to do in RL is push up the actions that make you better than average and push down the actions that make you worse than average. They call this advantage. There are multiple ways to compute it. I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it. At the end of the day, you want to reduce variance by trying to make this smaller, so it doesn’t magnify the variance of this one. That makes sense. This requires you to have a very good estimate of what average performance from a state would look like. This gets us back to the value function thing we were talking about earlier.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.
Named in this claim