Evidence receipt / belief
Published · transcript-backedPaul Christiano: belief
31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover
“What is a good explanation? And even when people are doing informal interpretability I think if you’re publishing in an ML conference and you want to say this is a good explanation, the way you would verify that would even if not like a formal set of causal intervention experiments.”
Source trail
Everything needed to verify it.
- Speaker
- Paul Christiano
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Even if the output is the same. Yeah. Or if the output or if it’s too expensive to check in this case. And to be clear, when we talk about formalizing, what is a good explanation? I think there is a little bit of work that pushes on this and it mostly takes this causal approach of saying, well, what should an explanation do? It should not only predict the output, it should predict how the output changes in response to changes in the internals. So that’s the most common approach to formalizing. What is a good explanation? And even when people are doing informal interpretability I think if you’re publishing in an ML conference and you want to say this is a good explanation, the way you would verify that would even if not like a formal set of causal intervention experiments. It would be some kind of ablation where then we messed with the inside of the model and it had the effect which we would expect based on our explanation. Anyways, back to the problems of mechanistic interpretability.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.