Evidence receipt / belief
Published · transcript-backedPaul Christiano: belief
31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover
“I think one of the things, one of the most important things for people concerned about misalignment to do is to ensure that adequate effort is put into this task.”
Source trail
Everything needed to verify it.
- Speaker
- Paul Christiano
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…How do you create the optimal conditions for it to want to be deceptive? Do you fine tune it on mindcomp or what are you doing? Yeah, so for deceptive alignment, I mean, I think it’s really complicated. Yeah, I think this is a very hard project. I think one of the things, one of the most important things for people concerned about misalignment to do is to ensure that adequate effort is put into this task. And we have, in fact, created good conditions in the laboratory and there’s like, real disagreement about how informative those will be. But okay, if you were to guess, for deceptive Lymen, how would you do this? An example is suppose you train an AI system that has some medium horizon goal. You train your AI system to, whatever, get some paperclips it wants to produce as many paperclips as it can over the next five days, just like select actions really aggressively for producing paperclips over the next five days. You do your Rlhf, you do your pre training, whatever. That’s like your phase one. You also ensure your AI system has a really good understanding of how it’s trained. So, like, this AI system wants paperclips and it understands everything about how it’s trained and everything about how it’s fine tuned. And you train on just like a lot of this data. And they say, okay, if we’ve done all that, we have this concern that if a system wants paperclips and understands really well how it’s trained, then if it’s going to be trained to get apples instead of paperclips, it’s just going to do some cost benefit and be like, really? While I’m being trained to get apples, I should do that. But I should do that whether or not even if I want paperclips, I should still do that. So training won’t really affect its values. It will just understand that it’s being trained to get apples. It will try and get apples, and then if you take it out of training, it will go back to getting paperclips. I think this exact setup has a number of problems or a number of disanalogies from the real world. But this is a setup which is relatively conducive to deceptive alignment like produce a system that wants one thing tell it a lot about its training, the kind of information you might expect a system would get, and then try and understand whether in fact it is able to then or it tends or sometimes under optimal conditions, in fact, continues pursuing paperclips, only pursuing apples when it thinks it’s being trained. All right, and this actually might be a good jumping off point into your current research. So in that case, where it wants to pursue paperclips, what would it look like for you to be say, well, we figured out why it wants paperclips, and we figured out how to make sure it doesn’t optimize for that at our expense when we deploy it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.