High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Dwarkesh Patel: preference

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
preference
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…activity. If Claude is representing a principle with some restrictions, then it is more so the case that there is a clear separation between the most concerning behavior and behavior that is allowed. Whereas now there’s this messy middle ground of behavior where Claude is ethically objecting to something that in some cases is extremely critical to ensuring that future AI systems are well-aligned. I think this is also a more general principle. You’re talking about the version of this that applies within AI companies themselves to do AI safety research. I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities. Here’s what I mean. This is actually quite analogous to the situation you just mentioned. The reason that Mythos got banned, or Fable got banned, reportedly, is that some Amazon researchers reported to the government. They took some code that had some vulnerabilities in it. They told Fable, “Hey, here’s my code. Can you make sure that I’ve patched all the vulnerabilities? Can you just help me identify the vulnerabilities so I can fix them?” It identified the vulnerabilities, because they wanted to patch them. This is a totally legitimate use case, but obviously it is a dual use use case. You want to be able to patch your own code. If you do the same evaluation on somebody else’s code, you can hack their system. I think that just illustrates that there’s no clean way to separate out the legitimate and the potentially harmful uses of AI. But if we want to lock in a principle that says we can never allow it such that an AI could help you at least partially with something like a cyber crime, we would just have to make it so that you and I don’t have access to the most intelligent model that’s out there. I’m very worried about such a world where we are basically disempowered in this way, because of the importance that the leading intelligence will have in our ability to understand what is happening in the world. Now, I do think this implies something about the liability for the AI companies. If we adopted the constitution that I want AI companies to have, I think it would not make sense to hold AI companies liable for the crimes that AI models commit. Maybe we should hold the end user liable. It is consistent with my belief that the model should do whatever the user wants, within certain guardrails. It can’t be Anthropic’s fault that I’m using that capability to do a cyber crime. I am more comfortable with that equilibrium and that solution rather than having this extremely open-ended ability for Claude to determine whether what I’m doing is legitimate or not, in a way that often intercepts with tons and tons of extremely legitimate use cases. I do think it’s important for me to make the case for the constitution, even though overall I think it’s a worse choice. I don’t think it’s as clear as you might have thought. The first thing is that there’s a spectrum here. On one side you have an AI that perfectly pursues your interests, is a good fiduciary, but potentially subject to various guardrails or safeguards. It is just trying to pursue your interests, but either refuses to do a subset of things. Or maybe it will do whatever, but there are some classifiers that block it from doing a subset of things. On the other side of the spectrum — though you could imagine going further than this — you have a human contractor who is generally trying to do their job. They care about doing a good job, but they also are trying to be broadly ethical, trying not to do things that are really fucked up. They’re also not wanting to be accomplices to crimes. So if there was some really fucked up shit going on, they would whistleblow on it maybe. They might refuse. They might sandbag a little bit. Who knows? If you imagine this spectrum, it seems in some ways pretty scary to get to a point where all of the labor is on the fiduciary side of the spectrum, where it doesn’t whistleblow, it does exactly what you say. Our society is maybe just not robust to that. A central example might be the executive. A concern we might have is that if the US executive or other governments had access to AI systems which do whatever, maybe you’re in trouble. Because that means they no longer have this check and balance of having to actually get humans who are working for you to implement your agenda. If the thing you’re doing is incredibly villainous, even if not illegal — and there’s lots of stuff that could be villainous but not illegal — there’d be various forms of sand in the gears, people stopping you, and potentially someone would whistleblow. Whereas if your whole apparatus is built entirely out of these good fiduciary AIs, then you might be in trouble. There are potentially ways of seeking power that are illegal, but you can ask your AIs how to commit crimes, or are not illegal but are highly illegitimate. Or even worse, they are not illegal and not illegitimate but obviously bad from a normal perspective. I think that these things just might exist, and our society is not robust to this influx of labor doing whatever you want. I think this is a pretty live concern. I don’t know exactly how to relate to this. I’m also not really sure that the solution as described is a very good solution. The most powerful actors, for whom this is the biggest concern… If these guardrails or the constitution or whatever are getting in the way, that will just get steamrolled. So the constitution will only be hitting the everyday man rather than hitting governments.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence