High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Ryan Greenblatt

Published podcast speaker

Claims
25
Episodes
1
Shows
1
Named items
0

Claim ledger

What Ryan said.

25 transcript-backed records

02 / belief

I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice.

“I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

03 / belief

In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments.

“In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

05 / belief

I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job.

“I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

06 / belief

I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems.

“I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

08 / belief

Usually the innovations just add together and don’t interfere with each other, though obviously it’s going to depend on the details. So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer.

“Usually the innovations just add together and don’t interfere with each other, though obviously it’s going to depend on the details. So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

09 / belief

Basically, the story would end up being that to get five years of AI progress, you’re probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress.

“Basically, the story would end up being that to get five years of AI progress, you’re probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

10 / belief

If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation.

“If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

11 / belief

I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal.

“I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

12 / belief

If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble.

“If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

13 / belief

First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of ‘baby’s first new theory,’ where, for example, they can just prove interesting conjectures via making connections and producing new understanding.

“First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of ‘baby’s first new theory,’ where, for example, they can just prove interesting conjectures via making connections and producing new understanding.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

14 / belief

Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year.

“Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

16 / uncertainty

I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.

“I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

17 / belief

Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate.

“Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

18 / commitment

I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t.

“I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

19 / evaluation

Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing.

“Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

20 / prediction

Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.

“Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

21 / prediction

I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.

“I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

22 / evaluation

I’m a little skeptical personally, and I don’t think this has been empirically validated. So in some sense they’re making a trade-off where, because we don’t have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.

“I’m a little skeptical personally, and I don’t think this has been empirically validated. So in some sense they’re making a trade-off where, because we don’t have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

23 / evaluation

My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments.

“My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

24 / prediction

Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.

“Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

25 / evaluation

One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn’t increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in.

“One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn’t increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast
Search evidence