Evidence receipt / belief
Published · transcript-backedTrenton Bricken: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…But that's also probably going to be true with whatever you find. You're still going to have variants that you haven't explained. Or you found a feature, but you don't know if it actually explains deception or something else instead. First of all, I'm not saying you shouldn't try the probing approach. We want to pursue the entire portfolio. We've got the therapist interrogating the patient by asking, "Do you have any troubling thoughts?" We've got the linear probe, which I'd analogize to a polygraph test where we're taking very high level summary statistics of the person's well-being. Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it. What percent of the alignment portfolio should mech interp be?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.