Evidence receipt / belief
Published · transcript-backedTrenton Bricken: belief
28 Mar 2024 Dwarkesh Podcast Sholto Douglas & Trenton Bricken — How LLMs actually think
“I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 28 Mar 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…What is a compelling explanation to you, especially for very smart models, of “I understand why it made this output and it was like for a legit reason.” If it's doing million line pull requests or something, what are you seeing at the end of that request where you're like, “yep good, that's chill.” So ideally you apply dictionary learning to the model. You've found features. Right now we're actively trying to get the same success for attention heads. You can do it for residual stream, MLP, and attention throughout the whole model. Hopefully at that point you can also identify broader circuits through the model that are more general reasoning abilities that will activate or not activate. But in your case where we're trying to figure out if this pull request should be approved or not. I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired. That would be an immediate thing. You can do more than that, but that would be an immediate one. But before I trace down on that, what does a reasoning circuit look like? What would that look like when you found it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.