Evidence receipt / evaluation
Published · transcript-backedDwarkesh Patel: evaluation
22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover
“I agree with the sentiment of obviously approaching this situation with caution, but I do want to point out the ways in which the analyses we've been using have been maximally adversarial.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 22 Aug 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah, I don't know. I'm hesitant to say it's like drugs for the model. Broadly speaking, I do basically agree that we have quite a lot of tools and options for training AIs, even AIs that are somewhat smarter than humans. I do think you have to actually do it. You had Eliezer on. I'm much more bullish on our ability to solve this problem, especially for AIs that are in what I think of as the "AI for AI safety sweet spot." This is a band of capability where they're sufficiently capable that they can be really useful for strengthening various factors in our civilization that can make us safe. That’s stuff like our alignment work, control, cybersecurity, general epistemics, maybe some coordination applications. There's a bunch of stuff you can do with AIs that, in principle, could differentially accelerate our security with respect to the sorts of considerations we're talking about. Let’s say you have AIs that are capable of that. You can successfully elicit that capability in a way that's not being sabotaged or messing with you in other ways. They can't yet take over the world or do some other really problematic form of power-seeking. If we were really committed, we could then go hard, put a ton of resources and really differentially direct this glut of AI productivity towards these security factors. We could hopefully control and understand, do a lot of these things you're talking about to make sure our AIs don't take over or mess with us in the meantime. We have a lot of tools there. You have to really try though. It's possible that those sorts of measures just don't happen, or they don't happen at the level of commitment, diligence, and seriousness that you would need. That’s especially true if things are moving really fast and there are other competitive pressures: “This is going to take compute to do these intensive experiments on the AIs. We could use that compute for experiments for the next scaling step.” There’s stuff like that. I'm not here saying this is impossible, especially for that band of AIs. It's just that you have to try really hard. I agree with the sentiment of obviously approaching this situation with caution, but I do want to point out the ways in which the analyses we've been using have been maximally adversarial. For example, let’s go back to the adult getting trained by Nazi children. Maybe the one thing I didn't mention is the difference in this situation, which is maybe what we're trying to get at with the drug metaphor. When you get an update, it's much more directly connected to your brain than a sort of reward or punishment a human gets. It's literally a gradient update down to the parameter of how much this would contribute to you putting this output rather than that output. Each different parameter we're going to adjust to the exact floating point number that calibrates it to the output we want. I just want to point out that we're coming into the situation pretty well. It does make sense, of course, if you're talking to somebody at a lab, to say, "Hey, really be careful." But for a general audience, should I be scared witless? You maybe should to the extent that you should be scared about things that do have a chance of happening. For example, you should be scared about nuclear war. But should you be scared in the sense of you’re doomed? No, you're coming up with an incredible amount of leverage on the AIs in terms of how they will interact with the world, how they're trained, and the default values they start with. I think it is the case that by the time we're building superintelligence, we'll have much better… Even right now—when you look at labs talking about how they're planning to align AIs—no one is saying we're going to just do RLHF. At the least, you're talking about scalable oversight. You have some hope about interpretability. You have automated red teaming. Hopefully, humans are doing a bunch more alignment work. I also personally am hopeful that we can successfully elicit from various AIs a ton of alignment work progress. There's a bunch of ways this can go. I'm not here to tell you 90% doom or anything like that. This is the basic reason for concern. Imagine that we're going to transition to a world in which we've created these beings that are just vastly more powerful than us. We've reached the point where our continued empowerment is just effectively dependent on their motives. It is this vulnerability to, “What do the AIs choose to do?” Do they choose to continue to empower us or do they choose to do something else?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.