High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Paul Christiano: belief

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“I think how you would automate interpretability if you wanted to right now is you take the process humans use it’s like great, we’re going to take that human process, train ML systems to do the pieces that humans do of that process and then just do a lot more of it.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
belief
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…The point you made about as you automate it is it because whatever work the automated alignment researcher is doing, you want to make sure you can verify it. I think it’s most of all a way you can automate. I think how you would automate interpretability if you wanted to right now is you take the process humans use it’s like great, we’re going to take that human process, train ML systems to do the pieces that humans do of that process and then just do a lot more of it. So I think that is great as long as your test decomposes into human sized pieces. And there’s just this fundamental question about large models which is like, do they decompose in some way into human sized pieces or is it just a really messy mess with interfaces that aren’t nice? And the more it’s the latter type, the harder it is to break it down to these pieces, which you can automate by copying what a human would do. And the more you need to say, okay, we need some approach which scales more structurally. But I think compared to most people, I am less worried about automating interpretability. I think if you have a thing which works that’s incredibly labor intensive, I’m fairly optimistic about our ability to automate it. Again, the stuff we’re doing, I think, is quite helpful in some worlds. But I do think the typical case like interpretability can add a lot of value. Without this, it makes sense what an explanation would mean in language like, this model is doing this because of whatever essay length thing. But you have trillions of parameters and you have all these uncountable number of operations. What does an explanation of why an output happened even mean?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence