High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Eric Jang: recommendation

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”

— Eric Jang

Source trail

Everything needed to verify it.

Speaker
Eric Jang
Attribution
Verified speaker
Claim type
recommendation
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…How would you design a baseline where it’s truly better? This is where in RL people use things like TD learning to better approximate the quality function, the Q that we mentioned earlier. You can try to subtract that from your return. Ideally, what you really want to do in RL is push up the actions that make you better than average and push down the actions that make you worse than average. They call this advantage. There are multiple ways to compute it. I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it. At the end of the day, you want to reduce variance by trying to make this smaller, so it doesn’t magnify the variance of this one. That makes sense. This requires you to have a very good estimate of what average performance from a state would look like. This gets us back to the value function thing we were talking about earlier.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Named in this claim

Books, apps, tools, and people.

Search evidence