High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Andrew Gordon Wilson: uncertainty

19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)

“I I sort of joke sometimes that if I hadn't met the airline passenger dataset, I don't know really what my life would be like now because it's it's driven so much of my research.”

— Andrew Gordon Wilson

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon Wilson
Attribution
Verified speaker
Claim type
uncertainty
Recorded
19 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…It was it was pretty linearly Some seasonality. And and you said here's 3 models. 1 is basically y equals MX plus c or something like that. Just a just a straight line. And I think was the second 1 something like 10 parameters and 10,000 parameters was the third 1. And almost everyone in the room said they preferred 1 or 2. And you said at the end of this conversation, I'm gonna convince you to prefer 3, which was 10,000 parameters. And I took a poll at the end and it did shift. And so that was promising. Yeah. I I sort of joke sometimes that if I hadn't met the airline passenger dataset, I don't know really what my life would be like now because it's it's driven so much of my research. And it's just amazing how people are biased towards choosing the linear function or the qubit polynomial even if in practice they're not making that choice. Like on CIFAR for example, it's not uncommon to use a neural net with tens of millions of parameters to fit a training set with tens of thousands of data points. And even before deep learning was popular, we were doing non parametric statistics where we were working with models like Gaussian processes that were inspired by taking infinite limits of neural nets that are more flexible than any neural net you can fit in memory. And other popular covariance functions as well like the RBF kernel really are like saying I wanna use an infinite order polynomial. And so even in these kind of classical statistical models, we're implicitly saying well if we're unhappy about the third choice, it's actually because it doesn't have enough parameters. We want infinitely many parameters, not just 10,000 parameters. And I think another way to say this is that parameter counting is a very bad proxy for model complexity. Right.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence