Evidence receipt / belief
Published · transcript-backedPaul Christiano: belief
31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover
“I think there’s a pretty good chance that you can just tell if someone is lying, like a brain emulation or an AI or whatever, unless they were aggressively selected.”
Source trail
Everything needed to verify it.
- Speaker
- Paul Christiano
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Do you think AI lie detectors will work where you kind of just look at the activations and not find explanations in the way you were talking about with Heuristics, but literally just like, here’s what truth looks like, here’s what lies look like. Let’s just segregate the lane space and see if we can identify the two. Yeah, I think to separate the like just train a classifier to do it is a little bit complicated for a few reasons and may not work. But if you just brought them to space and say like, hey, it’s like you want to know if someone’s lying, you get to interrogate them, but also you get to rewind them arbitrarily and make a million copies of them. I do think it’s pretty hard to lie successfully. You get to look at their brain even if you don’t quite understand what’s happening. You get to rewind them a million times. You get to run all those parallel copies into gradient descent or whatever. I think there’s a pretty good chance that you can just tell if someone is lying, like a brain emulation or an AI or whatever, unless they were aggressively selected. If it’s just they are trying to lie well rather than it’s like they were selected over many generations to be excellent at lying or something, then your ML system hopefully didn’t train it a bunch to lie. And you want to be careful about whether your training scheme effectively does that. Yeah, that seems like it’s more likely than not to succeed. And how possible do you think it will be for us to specify human verifiable rules for reasoning such that even if the AI is super intelligent, we can’t really understand why it does certain things. We know that the way in which it arises at these conclusions is valid. Like, if it’s trying to persuade us to something, we can be like, I don’t understand all the steps, but I know that this is something that’s valid and you’re not just making shit up.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.