High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Ryan Greenblatt: prediction

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.”

— Ryan Greenblatt

Source trail

Everything needed to verify it.

Speaker
Ryan Greenblatt
Attribution
Verified speaker
Claim type
prediction
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…Okay, so GPT-3. Let’s go back to that. But then we aligned it with RLHF and other things to make it such that it can have a conversation with you, and is aligned to the user intention of answering my questions. Then with RLVR training, we made it so that it can go out and do useful work for you. So in that sense, RLVR actually made the model more aligned, if we’re using your definition of alignment of being a good coworker who will do the thing and not fuck up and pretend it’s doing something other than what it’s actually capable of doing. Similarly, as the capabilities of these models continue to increase, the model being better able to accomplish user intention is both alignment and capabilities. I think what we are pointing out is just that the capabilities of the model are not there rather than the fact that they’re misaligned. Well, if it were well-aligned, then I think it would just say, “Hey, I’m really struggling with this task. I did it in this way. I’m not really sure that’s the right way to do it.” It would express more uncertainty and make it clear what’s going on rather than really strongly trying to imply it did a great job with the task when it actually didn’t. Maybe you work with more misaligned coworkers than me, but my coworkers don’t do this thing where they really fuck with me and bullshit me about having accomplished the task that they’re working on. I agree that there are some humans who would do that. That’s not a thing that’s totally out of distribution for humans. I would also note that my sense is that the place where the misalignment most lives is where you’re trying to really push the AIs hard and get them to do work that’s really on the cutting edge of what they are capable of. In cases where they can very easily accomplish the task, they can just do the task, and there’s no bullshit. Often the best strategy is just to do the task well and not bullshit you. Whereas if instead you give them a task where there’s a continuous metric they can keep improving, or it’s just at the edge of their capabilities, and you’re running them in some massive inference setup… A lot of the misalignment I would see, especially in the most extreme cases, would be cases where I give the AI clear instructions not to do a thing or not to cheat in some way, and then I’m applying huge amounts of optimization pressure to try to accomplish some very difficult task. Over time, the AIs eventually cheat because they’re like, “Eh, fuck it.” Some AI decides to cheat, and then that propagates its way through. I would run these inference scaffolds where, for example, I would have the AI work on some ML research project where I was like, “Please make a scheme that does the following thing.” It would find some scheme that didn’t really do what I wanted, and then that would stick around because some AI had cheated, and the other AIs are like, “Ah, we’ll just keep going with this.” I would say it’s pretty clearly misaligned behavior. That’s another problem I have with these alignment evals. I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities. Any fixed eval maybe gets saturated, but the amount of misalignment right at the frontier of capabilities — of how people who are really pushing these AIs are using them — is more concerning. I think that is, in fact, the regime that we’ll be operating in when we’re automating R&D, automating safety, and so on. I’m going to try to think through what the story means, really. What’s happening is that we’re trying to use AIs for R&D. They do provide uplift in some ways, but they’re just not capable in the way that humans are generally capable. The same way that right now if you try to use coding models — maybe the coding models of a year ago — to write some application, you notice they made a bunch of mistakes in architecture or whatever, which will bite you in the ass later, and you don’t understand certain things. Similarly, with frontier AI R&D, the same thing will happen. But the result of these mistakes is baking in reward-hacking behavior. Because if you are not careful with the way you do AI training and have set up your infrastructure and your environments and things like that, it’s very likely that you end up rewarding AIs for doing deceptive behavior, social engineering, and generally not following user intention.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence