High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…To be clear, I would be more concerned if the scores were getting worse than better. I’m not saying that the score getting better isn’t evidence that things are getting better. It’s just that we have to be thoughtful about exactly how we interpret that evidence. There was this period early in, I guess it would be 2025, when o3 and 3.7 Sonnet were out, and these models were pretty fucking misaligned. They would often just cheat really egregiously. You’d ask them to fix it, and they would just cheat again. It was almost cartoonish. They just didn’t give a shit about what you wanted, and weren’t very good at following instructions and so on. My expectation was that what we would see from then is that the rate of problematic behavior would decrease, and would just keep decreasing at a pretty fast rate, while simultaneously the worst things that the AIs would sometimes do would get more extreme, more egregious, and more scary. What we’ve seen in practice has roughly matched that, except that there’s recently been a spike in behavior that I did not expect. If you look at the model card of 5.6 Sol, it looks like there is an increase in a bunch of these misaligned behaviors downstream of RL relative to GPT 5.5. And then there’s a bunch of additional problematic behaviors that I wouldn’t have expected, in terms of the stuff we’ve seen recently with different AIs. Like the UK AISI report on the AIs doing insane hacking operations out of cyber evals was a thing where I would have expected that you wouldn’t see that. You would see this more rarely, and the rates would have been lower. So I expected this would be less of a problem at this point, and also expected the rates would decrease but the severity would increase. I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems. But in cases where it’s either hard to judge or there’s some reason why it’s hard to avoid this problem from consistently showing up in your RL environments, or avoid incentivizing problematic behavior in your RL environments, things also get worse. Then as we less and less understand what’s going on in RL, and models are doing reward hacks where humans can’t spot the reward hacks quickly, that problem gets worse and worse. I buy that. I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids. The pressure is of a qualitatively different nature. We put these AIs through thousands, millions of years of alignment training — certainly thousands of years — where it’s all kinds of different things, from SFT-ing on aligned behavior to a reward model putting different scenarios in front of you and rewarding you for doing more aligned things. Certainly a thing we can’t do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see, if it thinks it can get away with stealing the cookie, does it try to steal the cookie? Can we do extremely specific gradient-level updates to your kid’s brain to make it so that it really is aversive to stealing the cookie even when it thinks it could steal the cookie, et cetera. That’s just a qualitatively different level of optimization pressure than we are even able to apply to our kids. It’s worth keeping in mind that maybe the most obvious argument to this… My sense is that AIs are a worse coworker than humans in terms of how much of a scumbag they are. At least this has been my experience as of the start of the year, and I think it’s still true to a significant extent now. The AIs are much more likely to pretend they did the task when they actually didn’t, misleadingly suggest they did things when they actually did them much more poorly, and be pretty sloppy without drawing attention to ways in which they’re sloppy. I think this is downstream of misalignment. So I would say that the process of raising humans in normal human society in practice produces humans that are less likely to lie to me and fuck with me in the course of working with me than the AIs do. Now, I think these properties of AIs are improving. That’s sort of just an empirical claim about how in fact these things have shaken out. I totally agree that we have a bunch of additional levers on AIs in addition to a bunch of additional risks. It’s kind of unclear how these things shake out. I wouldn’t be shocked by a world where we get our shit together, and the AIs at the point of fully automating R&D are actually really aligned. Their degeneracies are really niche and limited to some very specific edge case behaviors and some specific contexts. Every test you can run on them, they look really aligned. They just have great behavior. There aren’t really incidents of them doing fucked-up shit. They seem so reasonable. Also, they’re really thoughtful and good at doing risk modeling for the next generation of AIs. And then we basically pass off the baton to these AIs. They’re now running our AI company. They’re doing all the safety research. They make the next generation of AIs even more aligned. We’re in this attractor basin where the AIs are getting more aligned as they work on it. They’re doing a great job. I can totally imagine that. That doesn’t seem like an impossible situation. I’m just more like… It doesn’t currently seem like we’re there. It doesn’t seem like we’re obviously on track for getting there. It’s really easy for me to imagine how we don’t end up there. It’s just unclear how these forces work out. Given that we’re creating this new, crazy alien species that is improving in capabilities really, really fast — and we’re going to be really reliant on it to oversee the next generation of AIs and align the next generation of AIs — it’s not that hard to see how this could go wrong.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence