High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Trenton Bricken: uncertainty

22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

“I don't know if they took predictions, they should have of like, "Hey, I'm going to fine tune ChatGPT on code vulnerabilities.”

— Trenton Bricken

Source trail

Everything needed to verify it.

Speaker
Trenton Bricken
Attribution
Verified speaker
Claim type
uncertainty
Recorded
22 May 2025
Publisher
Dwarkesh Podcast

Transcript context

…Coming back to the World War II question, you can think of it as a hierarchy of abstractions of trust here, where let's say you want to go and talk to Churchill. It helps a lot if you can verify that in that conversation, in that 10 minutes, he's being honest. This enables you to construct better meta narratives of what's going on. So maybe particle physics wouldn't help you there, but certainly the neuroscience of Churchill's brain would help you verify that he was being trustworthy in that conversation and that the soldiers on the front lines were being honest in their depiction of their description of what happened, this kind of stuff. So long as you can verify parts of the tree up, then that massively helps you build confidence. I think language models are also just really weird. With the emergent misalignment work. I don't know if they took predictions, they should have of like, "Hey, I'm going to fine tune ChatGPT on code vulnerabilities. Is it going to become a Nazi?" I think most people would've said no. That's what happened. How did they discover that it became a Nazi?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence