High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Dwarkesh Patel: uncertainty

22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover

“We're starting off with something that has this intricate representation of human values and it doesn't seem that hard to sort of lock it into a persona that we are comfortable with. I don't know what changes.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
uncertainty
Recorded
22 Aug 2024
Publisher
Dwarkesh Podcast

Transcript context

…For folks who are unfamiliar with the basic story, they might wonder, "Why would AI take over at all? What's the reason they would do that?" The general concern is that you're offering someone power, especially if you're offering it for free. Power, almost by definition, is useful for lots of values. We're talking about an AI that really has the opportunity to take control of things. Say some component of its values is focused on some outcome, like the world being a certain way, especially in a longer-term way such that its concern extends beyond the period that a takeover plan would encompass. The thought is that it's often the case that the world will be more the way you want it if you control everything, rather than if you remain an instrument of human will or some other actor, which is what we're hoping these AIs will be. That's a very specific scenario. If we're in a scenario where power is more distributed—especially where we're doing decently on alignment and we're giving the AI some amount of inhibition about doing different things, maybe we're succeeding in shaping their values somewhat—then it's just a much more complicated calculus. You have to ask, “What's the upside for the AI? What's the probability of success for this takeover path? How good is its alternative?” Maybe this is a good point to talk about how you expect the difficulties of alignment to change in the future. We're starting off with something that has this intricate representation of human values and it doesn't seem that hard to sort of lock it into a persona that we are comfortable with. I don't know what changes. Why is alignment hard in general? Let’s say we've got an AI. Let's bracket the question of exactly how capable it will be and talk about this extreme scenario where it really has the opportunity to take over. I think we might just want to avoid having to build an AI that we're comfortable with being in that position. But let's focus on it for simplicity's sake, and then we can relax the assumption. One issue is that you can't just test it. You can't give the AI this literal situation, have it take over and kill everyone, and then say, "Oops, update the weights." This is what Eliezer talks about. You care about its behavior in this specific scenario that you can’t test directly. We can talk about whether that's a problem, but that's one issue. There's a sense in which this has to be "off-distribution." You have to get some kind of generalization from training the AI on a bunch of other scenarios. Then there's the question of how it's going to generalize to the scenario where it really has this option.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence