High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Paul Christiano: belief

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
belief
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…Okay, got it. If it’s good to, like, slowing down AI is good, presumably because it gives you more time to do alignment. But alignment also helps speed up AI. Rlhf is alignment, and it help with Chat GPT, which sped up AI. So I actually don’t understand how the feedback loop nets out, other than the fact that if AI is happening, you need to do alignment at some point. Right? So, I mean, you can’t just not do alignment. Yes. I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment. You’re like, maybe alignment speeds up AI, but the only purpose of slowing down AI was to do it’s right. It could never come out ahead. I think the reason that you can come out ahead, the reason you could end up thinking the alignment was net negative, was because there’s a bunch of other stuff you’re doing that makes AI safer. Like, if you think the world is gradually coming better to terms with the impact of AI, or policies being made, or you’re getting increasingly prepared to handle the threat of authoritarian abuse of AI, if you think other stuff is happening that’s improving preparedness, then you have reason beyond alignment research to slow down AI. Actually. How big a factor is that? So let’s say right now we hit pause and you have ten years of no alignment, no capabilities, but just people get to talk about it for ten years. How much more does that prepare people than we only have one year versus we have no time is just dead time, where no research in alignment or capabilities happening.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence