High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Joe Carlsmith: belief

22 Aug 2024 Dwarkesh Podcast Joe Carlsmith — Preventing an AI takeover

“A model's verbal behavior—even if it has a planning process (which GPT-4, I think, doesn't in many cases)—doesn't necessarily reflect those criteria.”

— Joe Carlsmith

Source trail

Everything needed to verify it.

Speaker
Joe Carlsmith
Attribution
Verified speaker
Claim type
belief
Recorded
22 Aug 2024
Publisher
Dwarkesh Podcast

Transcript context

…Today I'm chatting with Joe Carlsmith. He's a philosopher and, in my opinion, a capital-G great philosopher. You can find his essays at joecarlsmith.com. So we have GPT-4, and it doesn't seem like a paperclipper thing. It understands human values. In fact, you can have it explain why being a paperclipper is bad or ask it to explain why the galaxy shouldn't be turned into paperclips. What has to happen such that eventually we have a system that takes over and converts the world into something valueless? When I'm thinking about misaligned AIs—or the type that I'm worried about—I'm thinking about AIs with a relatively specific set of properties related to agency, planning, awareness, and understanding of the world. One key aspect is the capacity to plan and make relatively sophisticated plans based on models of the world, where those plans are evaluated according to criteria. That planning capability needs to be driving the model's behavior. There are models that are, in some sense, capable of planning. But when they give output, it's not like that output was determined by some process of planning, like, “Here's what will happen if I give this output, and do I want that to happen?" The model needs to really understand the world. It needs to really be like, “Okay, here’s what will happen. Here I am. Here’s the politics of the situation.” It needs to have this kind of situational awareness to evaluate the consequences of different plans. Another thing to consider is the verbal behavior of these models. When I talk about a model's values, I'm referring to the criteria that end up determining which plans the model pursues. A model's verbal behavior—even if it has a planning process (which GPT-4, I think, doesn't in many cases)—doesn't necessarily reflect those criteria. We know that we're going to be able to get models to say what we want to hear. That's the magic of gradient descent. Modulo some difficulties with capabilities, you can get a model to output the behavior that you want. If it doesn't, then you crank it until it does. I think everyone admits that for suitably sophisticated models, they're going to have a very detailed understanding of human morality. The question is, what relationship is there between a model's verbal behavior—which you've essentially clamped, you're forcing the model to say certain things— and the criteria that end up influencing its choice between plans? I'm pretty cautious about assuming that when it says the thing I forced it to say—or when gradient descent has shaped it to say a certain thing— that that is a lot of evidence about how it's going to choose in a bunch of different scenarios. Even with humans, it's not necessarily the case that their verbal behavior reflects the actual factors that determine their choices. They can lie. They might not even know what they would do in a given situation, all sorts of stuff like that. It's interesting to think about this in the context of humans. There's that famous saying: "Be careful who you pretend to be, because you are who you pretend to be." You notice this with how culture shapes children. Parents will punish you if you start saying things that are inconsistent with your culture's values, and over time, you become like your parents, right? By default, it seems like it kind of works. Even with these models, it seems to work. They don’t really scheme against us. Why would this happen?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence