High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Trenton Bricken: uncertainty

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.”

— Trenton Bricken

Source trail

Everything needed to verify it.

Speaker
Trenton Bricken
Attribution
Verified speaker
Claim type
uncertainty
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…Yeah, it would be interesting. To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens. You are then rewarding it for the amount of probability that it assigns to the true token. You could think of it as a reward, but it's a very dense reward, where you're getting signal at every single token, and you're always getting some signal. Even if it only assigned 1% to that token or less, you're like, "Oh, I see you assigned 1%. Good job. Keep doing that." Yeah, upweight it.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence