Evidence receipt / belief
Published · transcript-backedDwarkesh Patel: belief
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“When I think about really smart people I know, they’re just not that effective in domains they don’t understand that well.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…Here are a few points. First, I bet if you look at randomly sampled training environments for Mythos, they’re actually very different from what it looks like to actually use the model in practice. My sense is that the RL distribution has really large deviations from the real-world data distribution, and it’s significantly smoothed over by a mix of transfer and having a small amount of data focused on the real world. My sense is that this will be a similar mechanism as how it works for the crazy, wildly, quite superhuman AI you get as a result of five years of AI progress on top of fully automated AI R&D. So let’s go through this a little bit. In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments. You build all these different RL environments where the AI has to adapt on the fly, learn on the fly, figure out what it should do, understand its situation better, and learn really quickly from feedback in order to succeed at its objective. And it has things like limited resources, and if it messes up, it can end up in a much worse position. If you train on a huge number of these environments, you will learn general skills of picking up context on the fly, and we’re already seeing this. It’s already the case that AIs are now much better at understanding roughly what’s going on and picking up context from a limited amount of information they’re given access to. Then those AIs could be put on the job at TSMC. Even though TSMC is not literally in their data distribution, their data distribution is really wide, and the AIs are extremely good on their data distribution, such that it transfers to picking up being good at being an engineer at TSMC and learning that on the fly. The way the AI gets good at being a TSMC engineer isn’t that it has a ton of cached knowledge on being a good TSMC engineer. It’s that it does the equivalent of some scaled-up version of in-context learning there. That’d be the most prosaic story. Obviously, there’s a bunch of different ways this could go. I think this maybe comes down to a difference of intuition about how far you can get. When I think about really smart people I know, they’re just not that effective in domains they don’t understand that well. But how long have they had to learn?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.