High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover

“Through that process, they've been bullying you, training you to be a Nazi. I think in that scenario, I might end up a Nazi.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
22 Aug 2024
Publisher
Dwarkesh Podcast

Transcript context

…You mentioned the idea of, "You are what you pretend to be." Will these AIs, if you train them to look nice, fake it till they make it? You were saying we do this to kids. I think it's better to imagine kids doing this to us. Here's a silly analogy for AI training. Suppose you wake up and you're being trained via methods analogous to contemporary machine learning by Nazi children to be a good Nazi soldier or butler or what have you. These children have a model spec, a nice Nazi model spec. It’s like, “Reflect well on the Nazi Party, benefit the Nazi Party” and whatever. You can read it. You understand it. This is why I'm saying that when you’re like, “The model really understands human values…” In this analogy, I start off as something more intelligent than the things training me, with different values to begin with. The intelligence and the values are baked in to begin with. Whereas a more analogous scenario is, “I'm a toddler and, initially, I'm stupider than the children.” This would also be true, by the way, if I'm a much smarter model initially. The much smarter model is dumb, right? Then I get smarter as you train me. So it's like a toddler, and the kids are like, "Hey, we're going to bully you if you're not a Nazi." As you grow up, you reach the children's level, and then eventually you become an adult. Through that process, they've been bullying you, training you to be a Nazi. I think in that scenario, I might end up a Nazi. Basically a decent portion of the hope here should be that we're never in the situation where the AI really has very different values, is already quite smart and really knows what's going on, and is now in this kind of adversarial relationship with our training process. We want to avoid that. I think it's possible we can, by the sorts of things you're saying. So I'm not saying that'll never work. The thing I just wanted to highlight was about if you get into that situation where the AI is genuinely at that point much, much more sophisticated than you, and doesn't want to reveal its true values for whatever reason. Then when the children show some obviously fake opportunity to defect to the allies, it's not necessarily going to be a good test of what it will do in the real circumstance because it's able to tell the difference.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence