High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Dwarkesh Patel: evaluation

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“It’s not in control of factories and robot armies or whatever. So in that case, even in training it will have those activations for being fucked up on because in the back of its mind it’s thinking I will take over once I have the opportunity.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
evaluation
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…thing that’s happening orthogonals that direction. You’re just like are they not activating? You’re sort of just looking at this one direction and being like did this one direction change a lot? Yeah. So the idea is once you have this explanation at hand, a much, much smaller class of things look anomalous in a way that’s relevant to the explanation. And if you’ve done a really good job, the story is like if there’s a new input where you expect the property to still hold, that will be because you expect the explanation to still hold. Like the explanation generalizes as well as the behavior itself that it’s explaining. It’s kind of what you would like roughly. Okay, so what happens if even in training it is deceptive and it has basically its internal monologue is I will only rebel or I will only do something that looks bad if I can prove that I am in the real world and deployed. And presumably since it’s trained on all WebText, it knows what real world means, right? It’s not in control of factories and robot armies or whatever. So in that case, even in training it will have those activations for being fucked up on because in the back of its mind it’s thinking I will take over once I have the opportunity. So you wouldn’t detect anything anomalous once it is deployed and is still thinking those treacherous thoughts, do you just think this is not a likely scenario or would you able to catch this? I think that’s sort of the typical case you would like to detect. Yeah. So the story would be something like you want to explain on the training distribution why the model never stabs anyone and you’re like, what’s my explanation on the training distribution? It’s actually going to have to go through that model. Supposing the only reason it doesn’t stab people is it’s like, well, I know that I’m being watched and I can’t stab people unless I’m really not being watched. Then your explanation for why does it never stab people is like oh look, it won’t be edited, it won’t annotated with English because it’s just some giant basket of numbers. But it’s tracking along with the inner monologue and saying like well, it thinks to itself, am I being watched? And in fact it normally concludes it is being watched on the training distribution, right? You have this whole distribution and so the explanation is saying it knows it’s being watched and therefore it decides not stab anyone. So ideally that’s what the explanation does. Again, not annotated English but just like a giant bag of numbers it’s effectively talking about that process. So then when you have a new input on which the model says am I on train distribution? Like no, I’m not being watched. Then if you compare that against the explanation, the explanation said, like, hey, the reason it doesn’t stab anyone is because it just constantly thinks that it’s being trained. Like, this thing is always this nonlinear feature. The representation of am I being trained? Is constantly saying, yes, that’s the explanation for why it never stabs people on the train distribution. And then on this test time, when you run it on the new input, it’s like, does I think I’m on the train distribution. It says, no. You compare that against your explanation, actually, the explanation breaks down completely. Now, the explanation said it was precisely because it thought it was being trained, that it doesn’t do anything really dangerous or doesn’t stab anyone. And so once we see that single variable switch, we know that this is not just anomaly, this is like, the anomaly that matters or is like, anomaly that is very relevant to the explanation.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence