Evidence receipt / evaluation
Published · transcript-backedDwarkesh Patel: evaluation
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies’ data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…There are a few things that are going on. One of the things is that over time we’re training AIs on increasingly complicated environments built by earlier AI systems, where humans don’t really fully understand what’s going on inside of these neural environments and don’t necessarily even roughly understand what’s going on with AI progress. So things are kind of drifting away from our understanding. We’re incentivizing all kinds of bad behaviors that we maybe even can’t notice. The AIs at some level understand these behaviors are bad, but the overall training process for those AIs also didn’t incentivize them to point out or fix these issues for us. Things are going off the rails. Also, when AIs are extremely, extremely capable, my view is that those AIs will be harder to align than current systems. For current systems, we have this feedback loop where basically we create an AI, we do some evaluations on it, we see that it has some kind of messed-up behavior that we can kind of quickly understand. Then we can go look in training and be like, “Oh, these training environments led to this problematic behavior. Let’s tweak that training data. Let’s introduce some additional training data to correct this other issue, and then move forward from there.” But in a regime where the AIs are extremely situationally aware, very, very capable, and we don’t necessarily understand what they’re doing, this feedback loop breaks down. I think it’s plausible that we’re going to see this behavioral feedback loop starting to break down over the next short period, as what AIs are already doing gets harder to understand. But I’m not sure about that. Okay, let’s break down both of those things one by one. As we can monitor them less and less, we have less ability to understand what they’re getting incentivized for. So even if it’s not the result of a malicious process… Let’s make it concrete for the audience. Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies’ data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized. If people are on Twitter, they will have seen all this stuff, but just to give people context. I think people will be aware of the OpenAI sandbox hack of the Hugging Face database. Something that has happened recently is when the UK AI Security Institute… Is everything getting relabeled “security” instead of “safety” these days? AI Security Institute, I think.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.