High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Dwarkesh Podcast / episode intelligence

Ryan Greenblatt – What happens once AI can automate AI research?

11 Aug 2026 40 published claims 2 attributable people

Speakers in the public record

Claim mix

belief 23uncertainty 6evaluation 5prediction 3commitment 2preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

40 published records

02 / belief

There’s not some sense in which the lawyer is really truly motivated by the good of the justice system. But I think the way current AIs are shaping up, certainly how Anthropic’s AI is shaping up, is with this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, help the user towards that end.

“There’s not some sense in which the lawyer is really truly motivated by the good of the justice system. But I think the way current AIs are shaping up, certainly how Anthropic’s AI is shaping up, is with this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, help the user towards that end.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

03 / belief

I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice.

“I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

05 / belief

In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments.

“In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

07 / belief

I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job.

“I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

08 / belief

I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems.

“I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

10 / belief

Usually the innovations just add together and don’t interfere with each other, though obviously it’s going to depend on the details. So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer.

“Usually the innovations just add together and don’t interfere with each other, though obviously it’s going to depend on the details. So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

12 / belief

Basically, the story would end up being that to get five years of AI progress, you’re probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress.

“Basically, the story would end up being that to get five years of AI progress, you’re probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

14 / belief

If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation.

“If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

15 / belief

What we’re basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and then also training the different data files going back from 2019 to 2026 with the current best algorithmic recipe. I think that will be interesting.

“What we’re basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and then also training the different data files going back from 2019 to 2026 with the current best algorithmic recipe. I think that will be interesting.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

16 / uncertainty

We just don’t have human-level robotics models yet. So you’re suggesting if we do that — if the AIs get really good at the verifiable stuff in chip design, et cetera, and then they get really good at building fabs — it’ll be the equivalent of going back to the 18th century and saying, “Okay, I don’t know what you guys are talking about in your parliament, but I’ve got a bunch of steamships and a bunch of Maxim guns.

“We just don’t have human-level robotics models yet. So you’re suggesting if we do that — if the AIs get really good at the verifiable stuff in chip design, et cetera, and then they get really good at building fabs — it’ll be the equivalent of going back to the 18th century and saying, “Okay, I don’t know what you guys are talking about in your parliament, but I’ve got a bunch of steamships and a bunch of Maxim guns.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

17 / belief

I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal.

“I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

19 / belief

If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble.

“If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

20 / belief

First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of ‘baby’s first new theory,’ where, for example, they can just prove interesting conjectures via making connections and producing new understanding.

“First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of ‘baby’s first new theory,’ where, for example, they can just prove interesting conjectures via making connections and producing new understanding.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

21 / belief

” There’s another quote that says, in part, and I’m taking it slightly out of context, “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude.

“” There’s another quote that says, in part, and I’m taking it slightly out of context, “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

22 / belief

Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year.

“Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

23 / uncertainty

As you were saying, by the time RLVR actually worked — even though you could have done it with less compute — we had to wait for oceans of compute, gigawatts of compute, to be available before people were doing this training, on the trajectory of compute continuing to increase so we make more breakthroughs. I don’t know.

“As you were saying, by the time RLVR actually worked — even though you could have done it with less compute — we had to wait for oceans of compute, gigawatts of compute, to be available before people were doing this training, on the trajectory of compute continuing to increase so we make more breakthroughs. I don’t know.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

25 / belief

I feel like they just kind of say vaguely pro-social things. It doesn’t feel like there’s necessarily a mind on the other end who’s like, “Okay, I have strictly evaluated the alignment situation right now, and I think we should stop,” rather than, “This is the kind of thing the AI companies would probably try to get the AIs to say.

“I feel like they just kind of say vaguely pro-social things. It doesn’t feel like there’s necessarily a mind on the other end who’s like, “Okay, I have strictly evaluated the alignment situation right now, and I think we should stop,” rather than, “This is the kind of thing the AI companies would probably try to get the AIs to say.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

27 / belief

I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids.

“I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

28 / uncertainty

I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.

“I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

29 / belief

Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate.

“Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

30 / commitment

I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t.

“I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

31 / evaluation

Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies’ data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized.

“Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies’ data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

32 / evaluation

Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing.

“Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

33 / prediction

Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.

“Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

34 / prediction

I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.

“I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

35 / commitment

I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can’t reason about given the fact that the training process is not public. So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training.

“I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can’t reason about given the fact that the training process is not public. So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

36 / evaluation

I’m a little skeptical personally, and I don’t think this has been empirically validated. So in some sense they’re making a trade-off where, because we don’t have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.

“I’m a little skeptical personally, and I don’t think this has been empirically validated. So in some sense they’re making a trade-off where, because we don’t have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

37 / preference

I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.

“I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

38 / evaluation

My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments.

“My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

39 / prediction

Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.

“Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

40 / evaluation

One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn’t increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in.

“One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn’t increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast
Search evidence