Evidence receipt / uncertainty
Published · transcript-backedTrenton Bricken: uncertainty
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah, it would be interesting. To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens. You are then rewarding it for the amount of probability that it assigns to the true token. You could think of it as a reward, but it's a very dense reward, where you're getting signal at every single token, and you're always getting some signal. Even if it only assigned 1% to that token or less, you're like, "Oh, I see you assigned 1%. Good job. Keep doing that." Yeah, upweight it.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.