High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Eric Jang: evaluation

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.”

— Eric Jang

Source trail

Everything needed to verify it.

Speaker
Eric Jang
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…We’re talking about how the good move, the out-of-distribution move, is a small fraction of all the moves that are played across all the games on which you’d want to train. This, of course, reminds me of how LLMs are trained with policy gradient methods. When Karpathy was on the podcast, he called it like “sucking supervision through a straw”. It’s interesting that this thing you’re saying—which would be intractable and prevents you from actually getting beyond a certain level in Go—is just by default how LLMs are trained? Right. This is not to say it doesn’t work. If you imagine increasing the number of games to millions of samples, you actually can get some meaningful supervision samples, so long as you find a way to mask out the supervision from these guys. This is where things start to get pretty related to RL in terms of advantage and baselines. Let’s look at the gradient variance of a very naive approach—I’m going to call it gradient RL—that’s basically the sum of rewards. The sum of rewards is the return. In our naive setup here, we only have an indicator variable for the return, where either you won or lost. In the case where you lost, your gradient is zero, so you don’t train on those examples, and when you won, you try to predict those things. You can think about this setup as a special case of this general formula here. The trouble here is that this is very high variance. When you multiply these terms out to compute the variance of the gradient, it’s equal to the expectation of… Just for simplicity, we can pretend this is on average zero or something if you’re centering it at no signal. The variance here basically means that you’re taking the square of this product term. You end up with a term that grows quadratically with T. When you have a setup like this, this thing acts as a coupling effect on top of these terms here. Let’s actually map this to an LLM case, and we can answer why LLMs only do one-step RL instead of a multi-step RL scenario. In LLMs, you have a decoder that might predict some words like “hello world”. In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens. In this case, I would say something like log(hel) + log(lo) + log(world). This is true, and if this term were one, then they would be the same thing. However, when you’re sampling, if you have a reward term assigned to every specific token, now you have these interaction effects between the cross-multiplication of these terms and these terms. The problem becomes, how do you ascribe the credit associated with every episode to all these different terms? The thing I’m confused on is what it would even look like to do it that way in LLMs, because you only get a reward at the end of the episode.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence