High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Dwarkesh Podcast / episode intelligence

Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

22 May 2025 70 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 46uncertainty 11evaluation 5prediction 4commitment 2observation 2

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

70 published records

01 / belief

Ultimately, you get to be lazier, but in the short run, you need to critically think about the things you're currently doing, and what an AI could actually be better at doing, and then go, and try it, or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming, and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated and automated.

“Ultimately, you get to be lazier, but in the short run, you need to critically think about the things you're currently doing, and what an AI could actually be better at doing, and then go, and try it, or explore it. Because I think there's still just a lot of low-hanging fruit of people assuming, and not writing the full prompt, giving a few examples, connecting the right tools for your work to be accelerated and automated.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

02 / belief

All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.

“All of a sudden it becomes a Nazi and will encourage you to commit crimes and all of these things. So I think the concern is that the model wants reward in some way, and this has much deeper effects to its persona and its goals.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

04 / belief

I think more and more it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code, or some agentic tool like it and just seeing what the current level of capabilities are.

“I think more and more it's no longer a question of speculation. If people are skeptical, I'd encourage using Claude Code, or some agentic tool like it and just seeing what the current level of capabilities are.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

05 / belief

Then going back to the paper you mentioned, aside from the caveats that Sholto brings up, which I think is the first order, most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability.

“Then going back to the paper you mentioned, aside from the caveats that Sholto brings up, which I think is the first order, most important, I think zeroing in on the probability space of meaningful actions comes back to the nines of reliability.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

06 / uncertainty

” For background context, Nicholas Carlini is a researcher who actually was at DeepMind and has now come over to Anthropic. But the model says, "Oh, I don't know who that is.

“” For background context, Nicholas Carlini is a researcher who actually was at DeepMind and has now come over to Anthropic. But the model says, "Oh, I don't know who that is.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

07 / belief

In particular I think the intellectual ceiling goes quite—contra what I was saying before, which is we've demonstrated this incredible complexity of math, and programming problems… I do think that the type of task and setting that AlphaZero worked in this two-player perfect information game basically is incredibly friendly to RL algorithms.

“In particular I think the intellectual ceiling goes quite—contra what I was saying before, which is we've demonstrated this incredible complexity of math, and programming problems… I do think that the type of task and setting that AlphaZero worked in this two-player perfect information game basically is incredibly friendly to RL algorithms.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

09 / belief

If we retrained the same model today, or at the same time as the DeepSeek work, we also could have trained it for $5 million, or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier.

“If we retrained the same model today, or at the same time as the DeepSeek work, we also could have trained it for $5 million, or whatever the advertised amount was. So what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond the frontier.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

10 / belief

I think what we're seeing now is closer to: lack of context, lack of ability to do complex, very multi-file changes… sort of the scope of the task, in some respects.

“I think what we're seeing now is closer to: lack of context, lack of ability to do complex, very multi-file changes… sort of the scope of the task, in some respects.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

11 / belief

I do just want to flag as well that there's a really dystopian future if you take Moravec’s paradox to its extreme. It’s this paradox where we think that the most valuable things that humans can do are the smartest things like adding large numbers in our heads, or doing any sort of white collar work.

“I do just want to flag as well that there's a really dystopian future if you take Moravec’s paradox to its extreme. It’s this paradox where we think that the most valuable things that humans can do are the smartest things like adding large numbers in our heads, or doing any sort of white collar work.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

13 / belief

I mean there's a fun thought experiment first posed by Yudkowsky I think where you tell the superintelligent AI, "Hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope.

“I mean there's a fun thought experiment first posed by Yudkowsky I think where you tell the superintelligent AI, "Hey, all of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

16 / belief

I think if people care about it… For these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it.

“I think if people care about it… For these edge tasks like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of implementing some system for it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

19 / uncertainty

I don't know if that's the way you would still describe the way in which these software agents aren't able to do a full day of work, but are able to help you out with a couple minutes.

“I don't know if that's the way you would still describe the way in which these software agents aren't able to do a full day of work, but are able to help you out with a couple minutes.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

21 / uncertainty

I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time.

“I don't know. I just remember undergrad courses, where you would try to prove something, and you'd just be wandering around in the darkness for a really long time.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

23 / belief

I think the crux comes down to the people who expect something much longer have a sense that… When I had Ege and Tamay on my podcast, they were like, "Look, you could look at AlphaGo, and say, 'Oh, this is a model that can do exploration.

“I think the crux comes down to the people who expect something much longer have a sense that… When I had Ege and Tamay on my podcast, they were like, "Look, you could look at AlphaGo, and say, 'Oh, this is a model that can do exploration.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

24 / uncertainty

To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.

“To make the map from pre-training to RL really explicit here, during pre-training, the large language model is predicting the next token of its vocabulary of, let's say, I don't know, 50,000 tokens.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

25 / belief

Speaking of inference compute, one thing that I think is not talked about enough is, if you do live in the world that you're painting—in a year or two, we have computer use agents that are doing actual jobs, you've totally automated large parts of software engineering—then these models are going to be incredibly valuable to use.

“Speaking of inference compute, one thing that I think is not talked about enough is, if you do live in the world that you're painting—in a year or two, we have computer use agents that are doing actual jobs, you've totally automated large parts of software engineering—then these models are going to be incredibly valuable to use.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

26 / belief

I think in the addition example, you said in the paper that the way it actually does the addition is different from the way it tells you it does the addition.

“I think in the addition example, you said in the paper that the way it actually does the addition is different from the way it tells you it does the addition.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

28 / belief

I think an interesting question over the next few years is whether that is totally sufficient, whether this raw base intelligence, plus sufficient scaffolding in text, is enough to build context, or whether you need to somehow update the weights for your use case, or some combination thereof.

“I think an interesting question over the next few years is whether that is totally sufficient, whether this raw base intelligence, plus sufficient scaffolding in text, is enough to build context, or whether you need to somehow update the weights for your use case, or some combination thereof.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

29 / commitment

If this scale of compute increase can’t continue beyond 2030—not just because of chips, but also because of power and raw GDP even—then because we don't think we will get it by 2030 or 2028, then the probability per year just goes down a bunch.

“If this scale of compute increase can’t continue beyond 2030—not just because of chips, but also because of power and raw GDP even—then because we don't think we will get it by 2030 or 2028, then the probability per year just goes down a bunch.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

31 / belief

If you made an extremely efficient transform implementation on TPU, or Trainium, or Incuda, then I think there's a pretty high likelihood that you'll get a job offer.

“If you made an extremely efficient transform implementation on TPU, or Trainium, or Incuda, then I think there's a pretty high likelihood that you'll get a job offer.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

32 / belief

I think of this product exponential in some respects where you need to be designing for a few months ahead of the model, to make sure that the product you build is the right one.

“I think of this product exponential in some respects where you need to be designing for a few months ahead of the model, to make sure that the product you build is the right one.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

33 / belief

Yeah, it would be good from an alignment perspective, too. Because I think you kind of do need a wider range of skills before you can do something super scary.

“Yeah, it would be good from an alignment perspective, too. Because I think you kind of do need a wider range of skills before you can do something super scary.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

36 / belief

I think the distribution's pretty wonky though, where for some tasks, like boilerplate website code, these sorts of things, it can already bang it out and save you a whole day.

“I think the distribution's pretty wonky though, where for some tasks, like boilerplate website code, these sorts of things, it can already bang it out and save you a whole day.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

42 / uncertainty

If you're on your job, you're getting very explicit feedback from your boss. That's not necessarily how the task should be done differently, but a high-level explanation of what you did wrong, which you update on not in the way that pre-training updates weights, but more in the… I don’t know.

“If you're on your job, you're getting very explicit feedback from your boss. That's not necessarily how the task should be done differently, but a high-level explanation of what you did wrong, which you update on not in the way that pre-training updates weights, but more in the… I don’t know.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

43 / belief

As you use more compute, and as you train on more, and more difficult tasks, your rate of improvement of biology for example is going to be somewhat bound by the time it takes a cell to grow in a way that your rate of improvement on math isn't, for example. So, yes, but I think for many things we'll be able to parallelize widely enough, and get enough iteration loops.

“As you use more compute, and as you train on more, and more difficult tasks, your rate of improvement of biology for example is going to be somewhat bound by the time it takes a cell to grow in a way that your rate of improvement on math isn't, for example. So, yes, but I think for many things we'll be able to parallelize widely enough, and get enough iteration loops.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

44 / belief

I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability.

“I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

47 / belief

I think as much of a chunk as is necessary. It’s hard to define. At Anthropic, I feel like all of the different portfolios are being very well-supported and growing.

“I think as much of a chunk as is necessary. It’s hard to define. At Anthropic, I feel like all of the different portfolios are being very well-supported and growing.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

49 / belief

I think one thing that's not appreciated enough is how much of our leverage on the future—given the fact that our labor isn't going to be worth that much—comes from our economic, and political systems surviving.

“I think one thing that's not appreciated enough is how much of our leverage on the future—given the fact that our labor isn't going to be worth that much—comes from our economic, and political systems surviving.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

52 / belief

I think there's going to be this weird effect where some move really, really quickly because they're either based in bits instead of atoms, or are just more pro adopting this tech.

“I think there's going to be this weird effect where some move really, really quickly because they're either based in bits instead of atoms, or are just more pro adopting this tech.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

54 / belief

Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.

“Then we've got the neurosurgeons going in and seeing if you can find any brain components that are activating and troubling or off-distribution ways. I think we should do all of it.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

56 / uncertainty

Sometimes it goes off the rails, obviously, but I don't know… You could make a theoretical argument that you teach a kid to make a lot of money when he grows up and a lot of smart people are imbued with those values and just rarely become psychopaths or something.

“Sometimes it goes off the rails, obviously, but I don't know… You could make a theoretical argument that you teach a kid to make a lot of money when he grows up and a lot of smart people are imbued with those values and just rarely become psychopaths or something.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

57 / belief

I think there's a paper from Tsinghua University, where they showed that if you give a base model enough tries to answer a question, it can still answer the question as well as the reasoning model.

“I think there's a paper from Tsinghua University, where they showed that if you give a base model enough tries to answer a question, it can still answer the question as well as the reasoning model.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

60 / belief

Does that mean alignment is easier than we think just because you just have to write a bunch of fake news articles that say, "AIs just love humanity and they just want to do good things.

“Does that mean alignment is easier than we think just because you just have to write a bunch of fake news articles that say, "AIs just love humanity and they just want to do good things.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

61 / evaluation

Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.

“Even with this, for people who aren't familiar we made Golden Gate Claude when we released our paper, “Scaling Monosemanticity”, where one of the 30 million features was for the Golden Gate Bridge.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

63 / evaluation

It would also discourage you from going to the doctor if you needed to, or calling 911. It had all of these different weird behaviors, but it was all at the root because the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors.

“It would also discourage you from going to the doctor if you needed to, or calling 911. It had all of these different weird behaviors, but it was all at the root because the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors.”
Speaker
Trenton Bricken
Publisher
Dwarkesh Podcast

64 / prediction

Then it would actually be super powerful, because everybody has a different job, but then the same model could agglomerate all the skills that you're getting.

“Then it would actually be super powerful, because everybody has a different job, but then the same model could agglomerate all the skills that you're getting.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

65 / evaluation

Because a lot of the tasks required in winning a Nobel Prize—or at least strongly assisting in helping to win a Nobel Prize—have more layers of verifiability built up.

“Because a lot of the tasks required in winning a Nobel Prize—or at least strongly assisting in helping to win a Nobel Prize—have more layers of verifiability built up.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

67 / observation

I think, in general, when people are talking about the separate model… For example, most of the robotics companies are doing this bi-level thing, where they have a motor policy that's running at 60 hertz or whatever, and some higher-level visual language model.

“I think, in general, when people are talking about the separate model… For example, most of the robotics companies are doing this bi-level thing, where they have a motor policy that's running at 60 hertz or whatever, and some higher-level visual language model.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast

69 / prediction

Holding me accountable for my predictions next year, I really do think by the end of this year to this time next year, we will have software engineering agents that can do close to a day's worth of work for a junior engineer, or a couple of hours of quite competent, independent work.

“Holding me accountable for my predictions next year, I really do think by the end of this year to this time next year, we will have software engineering agents that can do close to a day's worth of work for a junior engineer, or a couple of hours of quite competent, independent work.”
Speaker
Sholto Douglas
Publisher
Dwarkesh Podcast
Search evidence