High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Paul Christiano: belief

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it’s being trained to thinking it should defect.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
belief
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…All right, and this actually might be a good jumping off point into your current research. So in that case, where it wants to pursue paperclips, what would it look like for you to be say, well, we figured out why it wants paperclips, and we figured out how to make sure it doesn’t optimize for that at our expense when we deploy it? Yeah, I think that at a meta level in terms of what’s your protection like, I think what you want to be saying is, we have these examples in the lab of something bad happening. We’re concerned about the problem at all because we have examples in the lab. And again, this should all be an addition. I think you kind of want this defense in depth of saying, we also have this testing regime that would detect problems for the deployed model. We have our problems in the lab. We then have some techniques which we believe address these problems. We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it’s being trained to thinking it should defect. And then you can say on the lab, we have some understanding of when those techniques work and when they don’t. We have some understanding of the relevant parameters for the real system that’s deployed. And we have a. Reasonable margin of safety. So we have reasonable robustness on our story about when this works and when it doesn’t. And we can apply that margin of safety with a margin of safety to the real deployed system saying this is the kind of story you want to build towards in the long run. Do your best to produce all the failures you can in the lab or versions of them, do your best to understand what causes them, what kind of anomaly detection actually works for detecting this or what kind of filtering actually works and then apply that and that’s at the meta level. It’s not talking about what actually are those measures that would work effectively, which is obviously like what? I mean, a lot of alignment research is really based on this hypothetical of like, someday there will be AI systems that fail in this way. What would you want to do? Can we have the technologies ready either because we might never see signs of the problem or because we want to be able to move fast once we see signs of the problem. And obviously most of my life is in that I am really in that bucket. I mostly do alignment research. It’s just building out the techniques that do not have these failures such that they can be available as an alternative if in fact these failures occur. Got it. Okay.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence