Evidence receipt / belief
Published · transcript-backedRyan Greenblatt: belief
11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?
“I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice.”
Source trail
Everything needed to verify it.
- Speaker
- Ryan Greenblatt
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 11 Aug 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…I think there seems to be a crux here, which I think is just an empirical question we’ll see. How good is the transfer between getting really, really good at understanding the situation, getting up to speed, making progress over long periods in verifiable domains — which the AIs are obviously getting way, way better at really fast — to, “Okay, go talk to the president and convince him to do X thing.” Or, “You’re now in charge of Google. You must make Google a much more profitable company this quarter.” Let me try to spell out a few more arguments that are maybe relevant. One thing is, when looking at how the AIs have improved at essay writing… Let’s talk about that a little bit. You can get some data even on these domains. AIs will be able to get some data even on these domains when on a very fast progress trajectory. Maybe it’s hard to build a verifiable environment for “was your essay really good according to humans?” But you can do a bit of that. You can do some training. You can do some online training. The AIs will be able to do some online training based on real-world stuff. They’ll be able to have evals. They’ll be able to sample that. You can scale up the cadence at which you do this. The second thing is that in practice, when I just look at the transfer, it seems okay. I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice. Now, that doesn’t mean that Mythos is better than the best humans or something. It can still be significantly worse than typical human professionals at some aspect of their job while still being way better than GPT-4, which was not even close. So we’re talking about how much progress has come from data versus algorithmic progress over the last few years. That reminds me, I’m actually running an experiment with this with Jerry Han, who’s still a college student. What we’re basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and then also training the different data files going back from 2019 to 2026 with the current best algorithmic recipe. I think that will be interesting. I’m curious if you want to pre-register what amount of compute multipliers are coming from one versus the other.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.