Evidence receipt / evaluation
Published · transcript-backedAndrew Gordon Wilson: evaluation
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Trick question. Is predictive power the same as understanding? I agree with the idea that representation matters. Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems. And so if we're trying to build more general agents, we want to understand what sorts of representations are going to provide a better general description of the real world. And so that means we want to avoid things like shortcut learning and so on. If it's just gonna lead to good predictions and some contrived problem and not really in the real world. And so there's this question of like, can we understand what sorts of distribution shifts we might typically encounter? And can we build methods that have broadly more robustness to a variety of different types of realistic distributions? I think this connects to things like no free lunch sort of thinking. So the no free lunch theorems say that every model is equally good in expectation over all problems sampled uniformly from a distribution over all problems. And there are other no free lunch theorems that say no single learner can be good on all problems. The issue I think with these theorems is not their mathematical validity. What they're saying is correct under the assumptions they're making, but rather that the assumptions they're making are not a good description of the real world. So the real world is a small corner of all possible data sets. It's not drawn uniformly from a distribution of all possible problems. If we were to do that, we would mostly just get noise. And so the question then is to what extent is the structure across real world problems shared? And at what level of abstraction can we represent that shared structure? And so my contention is that the distribution over real world data is biased towards low comographic complexity. And so are some of the models that we started to to develop. Yeah. But can you give an example of where it was hard to confront a misconception? Why it mattered? And what the process involved?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.