High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Trenton Bricken: uncertainty

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time.”

— Trenton Bricken

Source trail

Everything needed to verify it.

Speaker
Trenton Bricken
Attribution
Verified speaker
Claim type
uncertainty
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…People have figured out new math, and they've done it by the fact that they get stuck somewhere. They're like, "Why am I getting stuck here? Let me think through this." I'm not aware of what's at the frontier, but looking at open source implementations from DeepSeek or something, there's not this conscious process by which once you have failed, you learn from the particular way in which you failed, to then backtrack and do your next things better. Just pure gradient descent, I wonder if that's a big limitation. I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time. Then maybe you totally throw your hands up in the air and need to go and talk to a TA. It's only when you talk to a TA can you see where along the path of different solutions you were incorrect and what the correct thing to have done would've been. That's in the case where you know what the final answer is, right? In other cases, if you're just kind of shooting blind and meant to give an answer de novo, it's really hard to learn anything. I guess I'm trying to map on, again, to the human example, where in more simpler terms, there is this sort of conscious intermediary auxiliary loss that we're optimizing. It's a very sort of self-conscious process. Forget about math. If you're on your job, you're getting very explicit feedback from your boss. That's not necessarily how the task should be done differently, but a high-level explanation of what you did wrong, which you update on not in the way that pre-training updates weights, but more in the… I don’t know.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence