Evidence receipt / belief
Published · transcript-backedJoe Carlsmith: belief
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“I think we might just want to avoid having to build an AI that we're comfortable with being in that position.”
Source trail
Everything needed to verify it.
- Speaker
- Joe Carlsmith
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…Maybe this is a good point to talk about how you expect the difficulties of alignment to change in the future. We're starting off with something that has this intricate representation of human values and it doesn't seem that hard to sort of lock it into a persona that we are comfortable with. I don't know what changes. Why is alignment hard in general? Let’s say we've got an AI. Let's bracket the question of exactly how capable it will be and talk about this extreme scenario where it really has the opportunity to take over. I think we might just want to avoid having to build an AI that we're comfortable with being in that position. But let's focus on it for simplicity's sake, and then we can relax the assumption. One issue is that you can't just test it. You can't give the AI this literal situation, have it take over and kill everyone, and then say, "Oops, update the weights." This is what Eliezer talks about. You care about its behavior in this specific scenario that you can’t test directly. We can talk about whether that's a problem, but that's one issue. There's a sense in which this has to be "off-distribution." You have to get some kind of generalization from training the AI on a bunch of other scenarios. Then there's the question of how it's going to generalize to the scenario where it really has this option. Is that even true? Because when you're training it, you can say, "Hey, here's a gradient update. If you get the takeover option on a platter, don't take it." And then, in red teaming situations where it thinks it has a takeover attempt, you train it not to take it. It could fail, but I feel like if you did this to a child, like "Don't beat up your siblings," the kid will generalize to, "If I'm an adult and I have a rifle, I'm not going to start shooting random people."…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.