High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Dwarkesh Patel: observation

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“However, the problem is you spend most of training in this regime, in the low pass rate regime.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
observation
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…What’s also tough here is that the distribution you’re sampling under is your policy’s distribution. If your policy has no chance of sampling “blue,” then you will never get a signal. Exactly. That’s modeled by the fact that your probability of sampling “blue” is extremely low. If you sample it, you do learn as much as you would have in supervised learning. In all other cases—99.99% of the time in an untrained model—you’re learning incredibly little from seeing that “halycon” or “told” is not the correct word. That’s what happens most of the time. You learn very little. If you put your pass rate on the X-axis and the bits you’re learning from a sample on the Y-axis, with 0%, 50%, and 100%, so at the end of training you’re here. If you have supervised learning, the negative log pass rate would look something like this. The entropy of a binary random variable would look like this, depending on whether you’re doing nats or bits. If you do bits it’s one here at the peak. This is like a coin flip. You learn the most from a coin flip. This is supervised learning. This is RL. However, the problem is you spend most of training in this regime, in the low pass rate regime. How fast you’re learning is a function of how many bits per sample you’re getting, and you’re getting very little signal here. If you chart the pass rate on a log scale—where at the beginning of training with a vocab size of 100K the pass rate is 1/100,000, then 1/10,000, 1/1,000, 1/100—what this graph looks like here, supervised learning would look like this, and then RL would look like that. And arguably you spend all your time here, potentially never even getting a single success. It’s a depressing plot in the sense that once you’re here, it’s not at all obvious how you get to there. Once you’re here you have something, but you actually spend all the time here in many RL problems. There’s a question of how you initialize so you’re at least at a non-zero pass rate. One more thing I’d like to add about bits per sample that’s very relevant to any machine learning problem is that there’s a connection to soft targets and distillation. If you have access to the logits, not just the one-hot token answer, if you have access to the soft targets the entropy of this distribution is far higher than the one-hot. There’s way more information in bits per sample in a soft label. That’s why distillation is so effective per sample. It’s giving you way more information per sample.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence