Evidence receipt / observation
Published · transcript-backedGrant Sanderson: observation
30 Jun 2026 Dwarkesh Podcast Grant Sanderson – AI and the future of math
“What’s interesting is that we started doing this over a year ago, and it’s fun to see a little bit of a tone shift in the way they talk about AI between mid-2025 and where we are now in 2026.”
Source trail
Everything needed to verify it.
- Speaker
- Grant Sanderson
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 30 Jun 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…Or at the very least, even if it couldn’t literally do every single thing white-collar humans can do, it would just have transformative effects in the way that getting gold in the IMO did not have transformative effects on the world. First of all, I do want to point out that I’m totally moving the goalpost here. When I interviewed Dario two or three years ago, I asked this question about why they haven’t been able to use their vast knowledge to connect ideas together and come up with a new discovery that way. That seems like the kind of thing where even a moderately intelligent person, if they knew this much information, would be able to come up with a medical diagnosis from the fact that this drug causes migraines, and this other thing does this, and maybe it’s the same drug that can cure both things. From an outsider’s perspective, mathematics seems clearly like a field where finding the counterexample to the unit distance problem conjecture was an example of this kind of thing. So it’s total goalpost moving. But then we can ask, what is the next benchmark? Now that AIs can do this thing we should have thought they’d be able to do, what is the next thing that would be quite impressive? There are a couple of candidate ideas here. One could be coming up with interesting problems in the first place, and the other is coming up with new kinds of objects or conceptualizations that create or unify fields. On the first one, right now we have these Millennium Prize problems because mathematicians have noted them. Riemann came up with this idea of the Riemann zeta function because he thought the zeros of this function would have some connection to the density of prime numbers. Figuring out why we think this is an interesting thing to study in the first place, why we are building this object and trying to answer questions about it—and answer this particular question about it—seems like the kind of thing that would be the next benchmark. You highlight two pretty good examples there. For anyone curious about the unit distance conjecture, there’s this really nice video by a math channel called Polylog where they talk about it. All of these discussions cause people to reflect on the process of doing math. They’re like, “Oh, this thing can do this impressive stuff. What does that mean for us?” One of the people in that video highlights this quote: “good mathematicians prove theorems, great mathematicians come up with conjectures, and the greatest mathematicians come up with definitions.” That’s more or less exactly your framing here. We need the conjecture generator and then the definition generator. That’s the premium-tier mathematician. I don’t understand how exactly you’d make that a benchmark. Usually, when I think of the word benchmark, I’m thinking of something that is a goalpost. The ball is through the goal or it’s not. You can clearly say, “Yes, this is done.” Partly that’s to be able to do things like RLVR, but also partly just to know that you haven’t moved the goalpost in answering. OpenAI can have their headline on disproving the unit distance conjecture because it’s a clear, distinct thing. It did it. Whereas imagine trying to have a headline on GPT-5.4 coming up with a really good conjecture. “We promise, everyone thinks it’s a good conjecture.” It just doesn’t land the same way. But maybe that doesn’t negate the fact that it’s the right thing to be thinking about. I would be surprised if it ever took the form of looking like a benchmark, where we have a score saying it’s passed because we can quantify how good a conjecture is. The nature of what it would take is probably that you’d feel a tone shift in conversations with mathematicians about the way it’s useful to work with. This series you referenced, which is not at all produced yet and probably won’t be for a couple of months, takes the form of us interviewing a lot of mathematicians. What’s interesting is that we started doing this over a year ago, and it’s fun to see a little bit of a tone shift in the way they talk about AI between mid-2025 and where we are now in 2026. In the real world, that’s a very short amount of time. In the AI world, that’s eons. We’re able to see this tone shift over those eons. I think the way you’d measure conjecture-generating ability is going to be more subjective, based on that tone shift. It will be mathematicians saying they’re not just using it to solve their problems, but that as they step back and decide what their research field should even be, a conversation with such-and-such model was genuinely helpful for that. I don’t think it’s likely you’d see it in the form of a headline saying this was yet another benchmark knocked down. It’s very interesting. The kinds of things you can’t make benchmarks for are also the kinds of things, at least in the current paradigm, you can’t easily train for. There’s really no fundamental difference between a benchmark and a training environment. It’s very easy to come up with some dichotomy of, “here’s a deep reason why AI can’t do a certain thing”, and then it turns out you’re just thinking about it the wrong way, and actually it can do it pretty soon thereafter. But I’m going to come up with—…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.