Evidence receipt / preference
Published · transcript-backedTerence Tao: preference
20 Mar 2026 Dwarkesh Podcast Terence Tao – Kepler, Newton, and the true nature of mathematical discovery
“When there are holes in the argument where none of the things are working, then what do you do? They can suggest random things, but often I find that trying to chase them down to make them work, and finding they don’t work, wastes more time than it saves.”
Source trail
Everything needed to verify it.
- Speaker
- Terence Tao
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 20 Mar 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…I feel like a big crux in these conversations about how good AI will be for science is, I think you said this, that they’re using existing techniques and modifying them. It would be interesting to understand how much progress one can make simply from using existing techniques. If I looked at the top math journals, how many of the papers are coming up with a new technique, whatever that means, versus using existing techniques on new problems? What is the overhang? If you just applied every known technique to every open problem, would that constitute a humongous uplift in our civilization’s knowledge, or would that not be that impressive and useful? This is a great question, and we don’t have the data to fully answer it yet. Certainly, a lot of work that human mathematicians do… When you take a new problem, one of the first things we do is we look at all the standard things that have worked on similar problems in the past, and we try them one by one. Sometimes that works, and that’s still worth publishing because the question was important. Sometimes they almost work, and you have to add one more wrinkle to it, and that’s also interesting. But the papers that go into the top journals are usually ones where the existing methods can kind of solve 80% of the problem, but then there is this 20% which is resistant and a new technique has to be invented to fill in the gaps. It’s very rare now that a problem gets solved with no reliance on past literature, where all the ideas come out of nowhere. That was more common in the past, but math is so mature now that it’s just so much of a handicap to not use the literature first. AI tools are getting really good at the first part of that, just trying all the standard techniques on a problem, often making fewer mistakes in applying them than humans. They still make mistakes, but I’ve tested these tools on little tasks that I can do, and sometimes they pick up errors that I make. Sometimes I pick up errors that they make. It’s about a tie right now. But I haven’t yet seen them take the next step. When there are holes in the argument where none of the things are working, then what do you do? They can suggest random things, but often I find that trying to chase them down to make them work, and finding they don’t work, wastes more time than it saves. I think some fraction of problems that we currently think are hard will fall from this method, especially the ones that haven’t received enough attention. With the Erdős problems, almost all of the 50 problems that were solved by AIs were ones for which there was basically no literature. Erdős posed the problem once or twice. Maybe some people tried it casually and couldn’t do it, but they never wrote up anything. But it turned out that there was a solution, and it was just combining this one obscure technique that not many people know about with some other result in the literature. That’s the median level of what AI can accomplish, and that’s really great. It clears out 50 of these problems. So I think you will see some isolated successes. But what we found… Some people have done large-scale sweeps of these Erdős problems. If you only focus on the success stories, the ones that get broadcast on social media, it looks amazing. All these problems that haven’t been solved for decades, now they’re falling. But whenever we do a systematic study, on any given problem an AI tool has a success rate of maybe 1% or 2%. It’s just that they can buy scale, and you just pick the winners. It looks great. I think there’ll be a similar thing happening with the hundreds of really prestigious, difficult math problems out there. s just that they can buy scale, and you just pick the winners. It looks great. I think there’ll be a similar thing happening with the hundreds of really prestigious, difficult math problems out there. Some AI may get lucky and actually solve them, and there will be some backdoor to solve the problem that everyone else missed. That will get a lot of publicity. But then people will try these fancy tools on their own favorite problem, and they will again experience the 1% to 2% success rate. There’ll be a lot of noise amongst the signal of when they’re working and when they’re not. It will be increasingly important to collect these really standardized datasets. There are efforts now to create a standard set of challenge problems for AIs to solve, and not just rely on the AI companies to only publish their wins and not disclose their negative results. That will maybe give more clarity as to where we’re actually at.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.