Evidence receipt / preference
Published · transcript-backedCristopher Moore: preference
4 Sept 2025 Machine Learning Street Talk The Day AI Solves My Puzzles Is The Day I Worry (Prof. Cristopher Moore)
“I'm just gonna go solve it anyway. And that's somehow because the real world presents us with examples of these problems where there is so much rich structure to sink your teeth into, whether that's the structure in text, the structure in images, and so on.”
Source trail
Everything needed to verify it.
- Speaker
- Cristopher Moore
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 4 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Very cool. Now, we're in the the regime of transformers, which are these huge over parameterized models that, you know, kind of predict the next token and sequence of tokens. And it's just so good to have you in the room with me because, you know, it's interesting to think about how they're limited in terms of, you know, learning and optimization, but also complexity and and computability perhaps as well in terms of the classes of of automata. You know, from from your expert position, how do you think about the limits of of these types of models? I mean, most of the work that I'm familiar with is where you can show that something is hard. But as some of your viewers know, traditionally in computer science, we say a problem is hard, we mean there exist hard examples, if those are cleverly designed by an adversary to be as hard as possible. And then in some interdisciplinary work at the boundary between statistical physics and machine learning and high dimensional statistics, different people, different names for it. There, you can prove that things are hard in the context of really random examples, so synthetic data which is drawn from some simple probabilistic model. And, of course, real world data is neither of these. Right? Real world data is not designed by an adversary to be as tricky as possible, and it's very far from random. It has all kinds of structure that both human intelligence and animal intelligence and artificial intelligence can exploit. And I think that's why a lot of people in machine learning, they often feel like, well, you know, proving something is hard in theory isn't really, you know, I don't don't care. I'm just gonna go solve it anyway. And that's somehow because the real world presents us with examples of these problems where there is so much rich structure to sink your teeth into, whether that's the structure in text, the structure in images, and so on. And I I think what's fascinating about the LLM transformer world is that I feel like a few years from now, we're going to look back and say, yeah, that architecture works. A lot of architectures work. Almost in some sense, any sufficiently rich architecture will work. What matters is that the world is structured and any architecture which is capable of capturing some of that structure is going to do well at prediction. And, you know, whether it does well at other things and the whole debate about whether they understand and so on, that's, you know, I have thoughts, but they're probably thoughts that other people have said just as well as I would or better. I do think, though, that some of this work on phase transitions is quite interesting. So this is where I've spent the past decade or 2, and, this is where some ideas from spin glass theory and the theory of disordered materials from physics has met with machine learning. And the idea here is that just as a magnet, which is if you heat a magnet up above a certain critical temperature, it suddenly loses its ability to hold a magnetic field. Below that temperature, the atoms will automatically align and you'll get a nice strong magnetic field. Above that, it just becomes very noisy. And there are similar phase transitions, in fact, using a lot of the same ideas from physics in machine learning. Mhmm. So if you have some ground truth and then some noise process, which then produces some noisy data, then depending on how much noise you have, that's a little bit like the temperature. learning. Mhmm. So if you have some ground truth and then some noise process, which then produces some noisy data, then depending on how much noise you have, that's a little bit like the temperature. If there's too much noise, then there's literally nothing you can do to discover the ground truth, the underlying pattern. It's just no longer present in the data. It's been washed out. If there's very little noise or if you like if the signal to noise ratio is very high, then it's very easy. And a lot of our favorite algorithms work very quickly, spectral algorithms, PCA, what have you, message passing algorithms, like belief propagation, and so on. Then there can also be these interesting middle ranges where you can find the ground truth if you do an exhaustive search. But we actually believe that there is no efficient algorithm that will succeed in that regime because you're wandering around in this high dimensional landscape of possible fits to the data, and the accurate ones are kind of hidden behind what in physics we call an energy barrier. And all of our favorite algorithms, whether they're Monte Carlo or gradient descent or message passing, get stuck for an exponential amount of time in a kind of amorphous mush of inaccurate fits to the data. And only if you have the luxury of exhaustive search would you find the accurate fit. So I love this work. I love its interdisciplinary nature. It connects with, like, replica theory and the stuff that Giorgio Parisi recently got the Nobel Prize for. I work with a number of his students and grand students. So it's a wonderful interdisciplinary community. That said, though, all of this is theory about random problems. And, again, real world problems have structure that can help guide us. And I think what's fascinating is that that real world structure seems very hard to mathematize. How do we talk about that structure? It's much more than just correlations. You know, the real world has all this rich hierarchy of objects and parts of objects. Ultimately, I feel like what LLMs are going to do and what transformers are going to do is help us mathematize that structure. I think that ultimately, we're going to learn a lot about the world from the fact that they succeed, in addition to learning things about them.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.