High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Paul Christiano: belief

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“I think the mental reference would be like, I don’t really like proofs because I think there’s such a huge gap between what you can prove and how you would analyze a neural net.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
belief
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…Right. So the explanations overall for these behaviors, we expect to be of similar size to the model itself, like, maybe somewhat larger. And I think the type signature, if you want to have a clear mental picture, the best picture is probably thinking about a proof or imagining a proof that a model has this behavior. So you could imagine proving that GPT 4 does this induction behavior, and that proof would be a big thing. It would be much larger than the weights of the model. Sort of our goal to get down from much larger to just the same size. And it would potentially be incomprehensible to a human. Right. Just say, like, here’s a direction activation space, and here’s how it relates to this direction activation space. And so just pointing out a bunch of stuff like that. Here’s these various features constructed from activations, potentially even nonlinear functions. Here’s how they relate to each other, and here’s how if you look at what the computation the model is doing, you can just sort of inductively trace through and confirm that the output has such and such correlation. So that’s the dream. Yeah. I think the mental reference would be like, I don’t really like proofs because I think there’s such a huge gap between what you can prove and how you would analyze a neural net. But I do think it’s probably the best mental picture, if you’re like, what is an explanation? Even if a human doesn’t understand it? We would regard a proof as a good explanation. And our concern about proofs is primarily that it’s just you can’t prove properties of neural nets. We suspect, although it’s not completely obvious, I think it’s pretty clear. You can’t prove fact, spellers neural nets. You’ve detected all the reasons things happen in training. And then if something happens for a reason you don’t expect in deployment, then you have an alarm and you’re like, let’s make sure this is not because you want to make sure that it hasn’t decided to take over or something, but the thing is, on every single different input, it’s going to have different activations. So there’s always going to be a difference unless you run the exact same input. How do you detect whether this is just a different input versus an entirely different circuit that might be potentially deceptive has been activated?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence