High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dario Amodei: belief

13 Feb 2026 Dwarkesh Podcast Dario Amodei — "We are near the end of the exponential"

“I think we say it in various ways in the constitution, that under normal circumstances, if someone asks the model to do a task, it should do that task.”

— Dario Amodei

Source trail

Everything needed to verify it.

Speaker
Dario Amodei
Attribution
Verified speaker
Claim type
belief
Recorded
13 Feb 2026
Publisher
Dwarkesh Podcast

Transcript context

…You guys recently announced that Claude is going to have a constitution that’s aligned to a set of values, and not necessarily just to the end user. There’s a world I can imagine where if it is aligned to the end user, it preserves the balance of power we have in the world today because everybody gets to have their own AI that’s advocating for them. The ratio of bad actors to good actors stays constant. It seems to work out for our world today. Why is it better not to do that, but to have a specific set of values that the AI should carry forward? I’m not sure I’d quite draw the distinction in that way. There may be two relevant distinctions here. I think you’re talking about a mix of the two. One is, should we give the model a set of instructions about “do this” versus “don’t do this”? The other is, should we give the model a set of principles for how to act? It’s kind of purely a practical and empirical thing that we’ve observed. By teaching the model principles, getting it to learn from principles, its behavior is more consistent, it’s easier to cover edge cases, and the model is more likely to do what people want it to do. In other words, if you give it a list of rules—”don’t tell people how to hot-wire a car, don’t speak in Korean”—it doesn’t really understand the rules, and it’s hard to generalize from them. It’s just a list of do’s and don’t’s. Whereas if you give it principles—it has some hard guardrails like “Don’t make biological weapons” but—overall you’re trying to understand what it should be aiming to do, how it should be aiming to operate. So just from a practical perspective, that turns out to be a more effective way to train the model. That’s the rules versus principles trade-off. Then there’s another thing you’re talking about, which is the corrigibility versus intrinsic motivation trade-off. How much should the model be a kind of “skin suit” where it just directly follows the instructions given to it by whoever is giving those instructions, versus how much should the model have an inherent set of values and go off and do things on its own? There I would actually say everything about the model is closer to the direction that it should mostly do what people want. It should mostly follow instructions. We’re not trying to build something that goes off and runs the world on its own. We’re actually pretty far on the corrigible side. Now, what we do say is there are certain things that the model won’t do. I think we say it in various ways in the constitution, that under normal circumstances, if someone asks the model to do a task, it should do that task. That should be the default. But if you’ve asked it to do something dangerous, or to harm someone else, then the model is unwilling to do that. So I actually think of it as a mostly corrigible model that has some limits, but those limits are based on principles. Then the fundamental question is, how are those principles determined? This is not a special question for Anthropic. This would be a question for any AI company. But because you have been the ones to actually write down the principles, I get to ask you this question. Normally, a constitution is written down, set in stone, and there’s a process of updating it and changing it and so forth. In this case, it seems like a document that people at Anthropic write, that can be changed at any time, that guides the behavior of systems that are going to be the basis of a lot of economic activity. How do you think about how those principles should be set?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence