Evidence receipt / belief
Published · transcript-backedPaul Christiano: belief
31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover
“If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn’t have any concise explanation for the distribution. So this first problem, just like it’s a whole different set of activations, I think you’re actually kind of okay, and then the thing that becomes more messy is like but the real world will not only be new samples of different activations, they will also be different in important ways.”
Source trail
Everything needed to verify it.
- Speaker
- Paul Christiano
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…You’ve detected all the reasons things happen in training. And then if something happens for a reason you don’t expect in deployment, then you have an alarm and you’re like, let’s make sure this is not because you want to make sure that it hasn’t decided to take over or something, but the thing is, on every single different input, it’s going to have different activations. So there’s always going to be a difference unless you run the exact same input. How do you detect whether this is just a different input versus an entirely different circuit that might be potentially deceptive has been activated? Yeah, I mean, to be clear, I think you probably wouldn’t be looking at a separate circuit, which is part of why it’s hard. You’d be looking at like, the model is always doing the same thing on every input. It’s always whatever it’s doing, it’s a single computation. So it’d be all the same circuits interacting in a surprising way. But yeah, this is just to emphasize your question even more. I think the easiest way to start is to just consider the IID case. So where you’re considering a bunch of samples, there’s no change in distribution. You just have a training set of like a trillion examples and then a new example from the same distribution. So in that case, it’s still the case. Every activation is different, but this is actually a very, very easy case to handle. So if you think about an explanation that generalizes across, like if you have a trillion data points and an explanation which is actually able to compress the trillion data points down to like, actually, it’s kind of a lot of compression. If you think about if you have a trillion parameter model and a trillion data points, we would like to find a trillion parameter explanation in some sense. So it’s actually quite compressed and sort of just in virtue of being so compressed, we expect it to automatically work essentially for new data points from the same distribution. If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn’t have any concise explanation for the distribution. So this first problem, just like it’s a whole different set of activations, I think you’re actually kind of okay, and then the thing that becomes more messy is like but the real world will not only be new samples of different activations, they will also be different in important ways. Like the whole concern was there’s these distributional shifts or like, not the whole concern, but most of the concern, maybe the point of having these explanations. I think every input is an anomaly in some ways, which is kind of the difficulty is if you have a weak notion of anomaly, any distribution shift can be flagged as an anomaly. And it’s like constantly getting anomalies. And so the hope of having such an explanation is to be able to say, like, here were the features that were relevant for this explanation or for this behavior. And a much smaller class of things are anomalies with respect to this explanation. Like, most anomalies wouldn’t change this. Like most ways you change your distribution won’t affect the validity of this explanation. For example, this explanation is saying models will tend to activate in the following direction. You don’t care about anything that’s happening orthogonals that direction. You’re just like are they not activating? You’re sort of just looking at this one direction and being like did this one direction change a lot? Yeah. thing that’s happening orthogonals that direction. You’re just like are they not activating? You’re sort of just looking at this one direction and being like did this one direction change a lot? Yeah. So the idea is once you have this explanation at hand, a much, much smaller class of things look anomalous in a way that’s relevant to the explanation. And if you’ve done a really good job, the story is like if there’s a new input where you expect the property to still hold, that will be because you expect the explanation to still hold. Like the explanation generalizes as well as the behavior itself that it’s explaining. It’s kind of what you would like roughly.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.