Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover
“Might be like Rlhf where we don’t know if it generalizes, but so far it makes your chat GPT thing better and you can also use it to make sure that chat GPT doesn’t tell you how to make a bioweapon.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 31 Oct 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah. So our quest is to design training methods for which we don’t expect them to lead to reward hacking or don’t expect them to lead to receptive alignment. Ideally that won’t be like a huge tax where people are like, well, we use those methods only if we’re really worried about reward hacking or receptive alignment. Ideally those methods would just work quite well and so people would be like, sure, I mean, they also address a bunch of other more mundane problems so why would we not use them? Which I think is like that’s sort of the good story. The good story is you develop methods that address a bunch of existing problems because they just are more principled ways to train AI systems that work better, people adopt them and then we are no longer worried about eg reward hacking or deceptive alignment. And to make this more concrete, tell me if this is the wrong way to paraphrase it, the example of something where it just makes a system better, so why not just use it, at least so far? Might be like Rlhf where we don’t know if it generalizes, but so far it makes your chat GPT thing better and you can also use it to make sure that chat GPT doesn’t tell you how to make a bioweapon. So yeah, it’s not a mixture of tax. Yeah. So I think this is right in the sense that using Rlhf is not really a tax. If you wanted to deploy a useful system, why would you not? Or it’s just very much worth the money of doing the training. RHF will address certain kinds of alignment failures that is like, where system just doesn’t understand or is changing. Next word, prediction. It’s like, this is the kind of context where human would do this wacky thing even though it’s not what we’d like. There’s like some very dumb alignment failures that will be addressed by it. But I think mostly the question is, is that true even for the sort of more challenging alignment failures that motivate concern in the field? I think RL doesn’t address most of the concerns that motivate people to be worried about alignment.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.