Evidence receipt / evaluation
Published · transcript-backedDwarkesh Patel: evaluation
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“With these models, at least so far, it doesn't seem to matter. They just get, "Hey, don't help people make bombs" or whatever, even if you ask in a different way how to make a bomb.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…You can hope that will work too, but there is a concern. I sometimes think about AI takeover scenarios via this spectrum of how much power we voluntarily transferred to the AIs. How much of our civilization did we hand to the AIs intentionally by the time they took over? Versus, how much did they take for themselves? Some of the scariest scenarios are where we have a really fast explosion to the point where there wasn't even a lot of integration of AI systems into the broader economy. But there's this really intensive amount of superintelligence concentrated in a single project or something like that. That's a quite scary scenario, partly because of the speed and people not having time to react. Then there are intermediate scenarios where some things got automated, maybe people handed the military over to the AIs or we have automated science. There are some rollouts and that’s giving the AIs power that they don't have to take. We're doing all our cybersecurity with AIs and stuff like that. Then there are worlds where you more fully transitioned to a kind of world run by AIs where, in some sense, humans voluntarily did that. Joe Carlsmith (cont’d) 00:19:20 Maybe there were competitive pressures, but you intentionally handed off huge portions of your civilization. At that point, it's likely that humans have a hard time understanding what's going on. A lot of stuff is happening very fast. The police are automated. The courts are automated. There's all sorts of stuff. Now, I tend to think a little less about those scenarios because I think they're correlated with being further down the line. Humans are hopefully not going to just say, "Oh yeah, you built an AI system, let's just..." When we look at technological adoption rates, it can go quite slow. Obviously there's going to be competitive pressures, but in general this category is somewhat safer. But even in this one, I think it's intense. If humans have really lost their epistemic grip on the world, they've handed off the world to these systems. Even if you're like, "Oh, there's laws, there's norms…" I really want us to have a really developed understanding of what's likely to happen in that circumstance, before we go for it. I get that we want to be worried about a scenario where it goes wrong. But again, what is the reason to think it might go wrong? In the human example, your kids are not maximally adversarial against your attempts to instill your culture on them. With these models, at least so far, it doesn't seem to matter. They just get, "Hey, don't help people make bombs" or whatever, even if you ask in a different way how to make a bomb. We're also getting better and better at this all the time. You're right in picking up on this assumption in the AI risk discourse of what we might call intense adversariality between agents that have somewhat different values. There's some sort of thought—and I think this is rooted in the discourse about the fragility of value and stuff like that—that if these agents are somewhat different, at least in the specific scenario of an AI takeoff, they end up in this intensely adversarial relationship. You're right to notice that's not how we are in the human world. We're very comfortable with a lot of different differences in values. A factor that is relevant is this notion that there are possibilities for intense concentration of power on the table. There is some kind of general concern, both with humans and AIs. If it's the case that there's some ring of power that someone can just grab that will give them huge amounts of power over everyone else, suddenly you might be more worried about differences in values at stake, because you're more worried about those other actors. We talked about this Nazi example where you imagine that you wake up and you're being trained by Nazis to become a Nazi. You're not right now. Is it plausible that we'd end up with a model that is in that sort of situation? As you said, maybe it's trained as a kid. It never ends up with values such that it's aware of some significant divergence between its values and the values that the humans intend for it to have. If it's in that scenario, would it want to avoid having its values modified? At least to me, it seems fairly plausible that the AI's values meet certain constraints. Do they care about consequences in the world? Do they anticipate that the AI's preserving its values will better conduce to those consequences? Then it's not that surprising if it prefers not to have its values modified by the training process.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.