Evidence receipt / prediction
Published · transcript-backedDwarkesh Patel: prediction
30 Jun 2026 Dwarkesh Podcast Grant Sanderson – AI and the future of math
“I think the more relevant thing is: what is the data on which whatever architecture or loss function you have is incentivizing you to produce?”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 30 Jun 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…th the world is that you receive a slip of paper, and someone says, “Can you predict what will come next?” You predict what will come next, and then your memory’s wiped. You get another slip of paper. Imagine that was done a whole bunch of times, and then what comes out on the other end. They say, “Look at this essay that you wrote.” You might look at that and say, “This is awful. That’s not the essay that I would’ve written.” The process of repeatedly predicting something is just pretty different from how you would think as a writer to compose it and think it through. In particular, what would probably happen is that you’re a slave to your context. You might be answering some question about a particular field, so you draw on all the context around that. But the connection where all the substance is going to come from is, by its nature, a very unlikely one. You can do all the RL that you want to try to get better in some way, but what’s the thing that’s specifically upweighting and incentivizing making these unlikely connections when the vast majority of them aren’t the predictable next token that would come in there? So it might be the case that you just have this intelligence locked inside that box, but it’s a weird way of interacting with it. The thing I’m curious about is: do you ever get any fruit by questioning the premise of how tokens are generated? I don’t think it would be as simple as manipulating the temperature, but are there any things that you can do that take the existing level of intelligence but find the right ways of sparking those connections that unlock these sorts of things that we’ve seen? Or do you just need a little bit more intelligence, such that at the level of prediction, it’s predicting that it should be making that lightning bolt to another field? I think it’s more productive to reason, instead of architecture or even loss function, about data. We have diffusion models that do text, and the kinds of things they produce are not of a wholly different character. They’ve just not been explored as much. I think the more relevant thing is: what is the data on which whatever architecture or loss function you have is incentivizing you to produce? It does seem like they’re getting better. Forget about math. We did have a couple of examples of this kind of thing, but if you just look at why they’re getting better at being autonomous agents… They’re in an environment where they’re autoregressively producing the step that says “Let’s step back and do a search over the whole codebase,” and then “Let’s step back and assess my mistake,” is the thing that works. I assume what happened in the case of progress in science or maybe in math is you have frontier math-like problems. Mathematicians have specifically designed them because they require connecting together two different fields. I’m guessing there’s all kinds of clever, partially synthetic ways to make harder and harder problems like that that require these kinds of connections—for example, by eliminating assumptions and still requiring the AI to get to the answer—and then it doesn’t really end up mattering what the loss function is. It’s really about, can you come up with an environment that incentivizes this ability? It feels like you should be able to. I certainly can’t speak to the correct ways of doing that to unlock all this, but it would just be pretty surprising. Don’t you think it would be surprising if, over the next three years, there weren’t a lot more of those lightning bolts?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.