High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Gordon Wilson: evaluation

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
evaluation
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…Trick question. Is predictive power the same as understanding? I agree with the idea that representation matters. Representation meaning sort of how you're solving the problem even if you're getting the same performance in a particular application. But the reason it matters is because different representations that are achieving the same performance might give you different performance than on different problems. And so if we're trying to build more general agents, we want to understand what sorts of representations are going to provide a better general description of the real world. And so that means we want to avoid things like shortcut learning and so on. If it's just gonna lead to good predictions and some contrived problem and not really in the real world. And so there's this question of like, can we understand what sorts of distribution shifts we might typically encounter? And can we build methods that have broadly more robustness to a variety of different types of realistic distributions? I think this connects to things like no free lunch sort of thinking. So the no free lunch theorems say that every model is equally good in expectation over all problems sampled uniformly from a distribution over all problems. And there are other no free lunch theorems that say no single learner can be good on all problems. The issue I think with these theorems is not their mathematical validity. What they're saying is correct under the assumptions they're making, but rather that the assumptions they're making are not a good description of the real world. So the real world is a small corner of all possible data sets. It's not drawn uniformly from a distribution of all possible problems. If we were to do that, we would mostly just get noise. And so the question then is to what extent is the structure across real world problems shared? And at what level of abstraction can we represent that shared structure? And so my contention is that the distribution over real world data is biased towards low comographic complexity. And so are some of the models that we started to to develop. Yeah. But can you give an example of where it was hard to confront a misconception? Why it mattered? And what the process involved?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence