High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ryan Greenblatt: belief

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems.”

— Ryan Greenblatt

Source trail

Everything needed to verify it.

Speaker
Ryan Greenblatt
Attribution
Verified speaker
Claim type
belief
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…But how do we falsify this? Because it seems like this prediction of doom is basically saying that as things look better and better empirically, things will actually be worse and worse for our ability to not get taken over. To be clear, I would be more concerned if the scores were getting worse than better. I’m not saying that the score getting better isn’t evidence that things are getting better. It’s just that we have to be thoughtful about exactly how we interpret that evidence. There was this period early in, I guess it would be 2025, when o3 and 3.7 Sonnet were out, and these models were pretty fucking misaligned. They would often just cheat really egregiously. You’d ask them to fix it, and they would just cheat again. It was almost cartoonish. They just didn’t give a shit about what you wanted, and weren’t very good at following instructions and so on. My expectation was that what we would see from then is that the rate of problematic behavior would decrease, and would just keep decreasing at a pretty fast rate, while simultaneously the worst things that the AIs would sometimes do would get more extreme, more egregious, and more scary. What we’ve seen in practice has roughly matched that, except that there’s recently been a spike in behavior that I did not expect. If you look at the model card of 5.6 Sol, it looks like there is an increase in a bunch of these misaligned behaviors downstream of RL relative to GPT 5.5. And then there’s a bunch of additional problematic behaviors that I wouldn’t have expected, in terms of the stuff we’ve seen recently with different AIs. Like the UK AISI report on the AIs doing insane hacking operations out of cyber evals was a thing where I would have expected that you wouldn’t see that. You would see this more rarely, and the rates would have been lower. So I expected this would be less of a problem at this point, and also expected the rates would decrease but the severity would increase. I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems. But in cases where it’s either hard to judge or there’s some reason why it’s hard to avoid this problem from consistently showing up in your RL environments, or avoid incentivizing problematic behavior in your RL environments, things also get worse. Then as we less and less understand what’s going on in RL, and models are doing reward hacks where humans can’t spot the reward hacks quickly, that problem gets worse and worse. I buy that. I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids. The pressure is of a qualitatively different nature. We put these AIs through thousands, millions of years of alignment training — certainly thousands of years — where it’s all kinds of different things, from SFT-ing on aligned behavior to a reward model putting different scenarios in front of you and rewarding you for doing more aligned things. Certainly a thing we can’t do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see, if it thinks it can get away with stealing the cookie, does it try to steal the cookie? Can we do extremely specific gradient-level updates to your kid’s brain to make it so that it really is aversive to stealing the cookie even when it thinks it could steal the cookie, et cetera. That’s just a qualitatively different level of optimization pressure than we are even able to apply to our kids.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence