Evidence receipt / evaluation
Published · transcript-backedTerence Tao: evaluation
20 Mar 2026 Dwarkesh Podcast Terence Tao – Kepler, Newton, and the true nature of mathematical discovery
“He had some preconceived theories first. It seems like this is less and less the way we make progress, just because the data is so much more massive and useful.”
Source trail
Everything needed to verify it.
- Speaker
- Terence Tao
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 20 Mar 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…The take I want to try on you is that Kepler was a high-temperature LLM. Newton comes up with this explanation of why the three laws of planetary motion must be true. Of course, the way that Kepler discovers the laws of planetary motion, or figures out the relative orbits of the different planets, is as you say a work of genius. But through his career, he’s just trying random relationships. In fact, in the book in which he writes down the third law of planetary motion, it’s an aside on The Harmonics of the World, which is just a book about how all these different planets have these different harmonies. And the reason there’s so much famine and misery on Earth is because the Earth is mi-fa-mi, that’s the note of Earth. It’s all this random astrology, but in there is the cube-square law, which tells you what relationship the period has to a planet’s distance from the Sun. As you were detailing, if you add that to Newton’s F=ma and the equation for centripetal acceleration, you get the inverse-square law. And so Newton works that out. But the reason I think this is an interesting story is that I feel LLMs can do the kind of thing of trying random relationships for twenty years, some of which make no sense, as long as there’s a verifiable data bank like Brahe’s dataset. “Ok, I’m going to try out random things about musical notes, Platonic objects, or different geometries, I have this bias that there’s some important thing about the geometry of these orbits.” Then one thing works. As long as you can verify it, these empirical regularities can then drive actual deep scientific progress. Traditionally, when we talk about the history of science, idea generation has always been the prestige part of science. A scientific problem comes with many steps. You have to identify a problem, and then you have to identify a good, fruitful problem to work on. Then you need to collect data, figure out a strategy to analyze the data, and make a hypothesis. At this point, you need to propose a good hypothesis, and then you need to validate. Then you need to write things up and explain. There are a dozen different components. The ones we celebrate are these eureka genius moments of idea generation. Kepler certainly had to cycle through many ideas, several of which didn’t work. I bet there were many that he didn’t even publish at all because they just didn’t fit. That’s an important part of the process, trying all kinds of random things and seeing if they worked. But as you say, it has to be matched by an equal amount of verification, otherwise it’s slop. We celebrate Kepler, but we should also celebrate Brahe for his assiduous data collection, which was ten times more precise than any previous observation. That extra decimal point of accuracy was essential for Kepler to get his results. He was using Euclidean geometry and the most advanced mathematics he could use at the time to match his models with the data. All aspects had to be in play: the data, the theory, and the hypothesis generation. I’m not sure nowadays that hypothesis generation is the bottleneck anymore. Science has changed in the century since. Classically, the two big paradigms for science were theory and experiment. Then in the 20th century, numerical simulation came along, so you can do computer simulations to test theories. Finally, in the late 20th century, we had big data. We had the era of data analysis. A lot of new progress is actually driven now by analyzing massive datasets first. You collect large datasets and then draw patterns from them to deduce thoughts. This is a little bit different from how science used to work, where you make a few observations or have one out-of-the-blue idea, and then collect data to test your idea. That’s the classic scientific method. Now it’s almost reversed. You collect big data first, and then you try to get hypotheses from it. Kepler was maybe one of the first early data scientists, but even he didn’t start with Tycho’s dataset and then analyze it. He had some preconceived theories first. It seems like this is less and less the way we make progress, just because the data is so much more massive and useful. Oh, interesting. I feel like the 20th-century science that you’re describing actually very well describes what happened with Kepler. He did have these ideas—1595 and ‘96 is where he comes up with the polygons and then the Platonic objects theory—but they were wrong. Then a few years later, he gets Brahe’s data, and it’s only after twenty years of trying random things that he gets this empirical regularity. It actually feels a bit closer to Brahe’s data being analogous to some massive data bank of simulations, and now that you’ve got the data, you can keep trying random things. If it wasn’t for that, Kepler would be out there just writing books about harmonics and Platonic objects, and there would be nothing to actually verify against.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.