Evidence receipt / belief
Published · transcript-backedTrenton Bricken: belief
28 Mar 2024 Dwarkesh Podcast Sholto Douglas & Trenton Bricken — How LLMs actually think
“I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first.”
Source trail
Everything needed to verify it.
- Speaker
- Trenton Bricken
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 28 Mar 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…Okay, so let's go back to GPT-7. First of all, is this sort of like a linear tax on any model to figure it out? Even before that, is this a one time thing you had to do or is this the kind of thing you have to do on every output? Or just one time it's not deceptive and we're good to roll? So you do dictionary learning after you've trained your model and you feed it a ton of inputs and you get the activations from those. Then you do this projection into the higher dimensional space. So the method is unsupervised in that it's trying to learn these sparse features. You're not telling them in advance what they should be but, it is constrained by the inputs you're giving the model. Two caveats here. One, we can try and choose what inputs we want. So if we're looking for theory of mind features that might lead to deception, we can put in the sycophancy data set. Hopefully at some point we can move into looking at the weights of the model alone, or at least using that information to do dictionary learning. I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first. So what's the cost of this? Can you repeat the last sentence? About the weights of the model alone.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.