High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Eric Jang

Published podcast speaker

Claims
20
Episodes
1
Shows
1
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

paper / recommends

Generalized Advantage Estimation

“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”

Dwarkesh Podcast · 15 May 2026

Evidence receipt · Source ↗

Claim ledger

What Eric said.

7 transcript-backed records

01 / evaluation

There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.

“There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

03 / evaluation

You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.

“You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

05 / evaluation

In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.

“In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast

06 / evaluation

So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.

“So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.”
Speaker
Eric Jang
Publisher
Dwarkesh Podcast
Search evidence