Evidence receipt / belief
Published · transcript-backedPaul Christiano: belief
31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover
“I think as you move beyond that, it really depends how you’re deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weights sitting there, I think you mostly have addressed the risk from the weights just sitting there.”
Source trail
Everything needed to verify it.
- Speaker
- Paul Christiano
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Sorry if you’re already getting to it, but the part I’m most curious about is separate from the ways in which other people might fuck with it, it’s isolated. What is the point at which we satisfied? It in and of itself is not going to pose a risk to humanity. It’s human level, but we’re happy with it. Yeah. So I think here I listed maybe the two most simple ones that start out like security. Internal controls, I think become relevant immediately and are very clear why you care about them. I think as you move beyond that, it really depends how you’re deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weights sitting there, I think you mostly have addressed the risk from the weights just sitting there. Now, what you’re talking about for risk is mostly, and maybe there’s some blurriness here of how much internal controls captures not only employees using the model, but anything a model can do internally. You would really like to be in a situation where your internal controls are robust not just to humans but to models potentially like E-G-A model shouldn’t be able to subvert these measures and you care just as you care about are your measures robust if humans are behaving maliciously? You care about are your measures robust if models are behaving maliciously? So I think beyond that if you’ve then managed the risk of just having the weight sitting around. Now we talk about in some sense most of the risk comes from doing things with the model. You need all the rest so that you have any possibility of applying the brakes or implementing a policy. But at some point as the model gets competent you’re saying like okay, could this cause a lot of harm? Not because it leaks or something, but because we’re just giving it a bunch of actuators. We’re deploying it as a product and people could do crazy stuff with it. So if we’re talking not only about a powerful model but like a really broad deployment of just something similar to the Open Eyes API where people can do whatever they want with this model and maybe the economic impact is very large. So in fact, if you deploy that system it will be used in a lot of places such that if AI systems wanted to cause trouble it would be very very easy for them to cause catastrophic harms. Then I think you really need to have some kind of I mean, I think probably the science and discussion has to improve before this becomes that realistic. But you really want to have some kind of alignment analysis, guarantee of alignment before you’re comfortable with this. And so by that I mean you want to be able to bound the probability that someday all the AI systems will do something really harmful. That there’s some thing that could happen in the world that would cause these large scale correlated failures of your AIS. And so for that there’s sort of two categories that’s like one, the other thing you need is protection against misuse of various kinds which is also quite hard. And by the way, which one are you worried about more misuse or misalignment?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.