Evidence receipt / evaluation
Published · transcript-backedAndrew Gordon Wilson: evaluation
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“Like, once a certain number of people believe something, it's very, very hard to change their minds no matter what you say. And I think as a consequence, we've been in all sorts of local minima in machine learning and AI research because we haven't been able to get unstuck from these erroneous beliefs.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Folks, that interview with Andrew was absolutely amazing. Keith came over to The UK and we did it in my home studio together a few weeks ago. It's been on Patreon for a little while and I updated so much based on that interview. Andrew was absolutely brilliant. So I know you're gonna love it. But before we kick off, you've probably heard that human data is kind of the dirty secret of Silicon Valley. You know, human data is the reason why these AI models work so well. Because the open AIs and the Anthropics, what they do is they hire humans to do things like data securation and evaluation and post training. And there is a ridiculous kind of uplift from using this human data. Our sponsor, Prolific, what they wanna do is produce the first survey on how human data is being used in AI. And if you volunteer and fill out this form for them, they will give you first access to see how you compare. So, I'd really appreciate it if you did that. There's no personally identifiable information. Link is in the description. And we are also sponsored by Twofer AI Labs. They are an incredible research lab based in Zurich. They've just upgraded their office. They've got an amazing new office. They've hired 13 research engineers in the last year, doing things like reasoning and the ARC challenge. You've probably seen some of the papers that they've published on that. But they have ambitions to build their own foundation models from scratch. They've got an amazing culture. And Benjamin Crusier, the the director, is also very interested in AI safety. So he's going through the Yudkowsky book at the moment. So if that seems like a fit for you, please get in touch with Benjamin Crusier. Go to 2forlabs.ai or look in the description. And also, MLST is sponsored by Cyber Fund. Enjoy the show, folks. Well, Andrew, much of your work challenges conventional wisdom. Is that hard to do? Is there resistance in challenging strongly held beliefs? So, yes. I mean, this is what happens when you challenge conventional wisdom. But in some sense, I think that we should always be trying to do that because otherwise, we're just preaching to the choir and then what's the point? But if no 1 if you're not changing anyone's beliefs about anything, then maybe it doesn't make a difference. And so I think it's important to really try to understand what do a lot of people believe that might be wrong and then just unpack that. And it's also very exciting and fun, but it's challenging because of course, the initial instinct will be to resist whatever you're saying. But then over time, and if you try hard enough and if you talk to amazing communicators like you 2, then you can start to have an influence. And I think that's really important because so much progress has been stalled, think, by just getting stuck on misconceptions. Like, once a certain number of people believe something, it's very, very hard to change their minds no matter what you say. And I think as a consequence, we've been in all sorts of local minima in machine learning and AI research because we haven't been able to get unstuck from these erroneous beliefs. And there's a whole roster of things like this, like the role of implicit biases of stochastic optimization and generalization, I think is significant, but also significantly overstated. How we can have really large models that also generalize well even when there's a small number of data points is something that is not very well recognized. And in fact, I think is 1 of the primary drivers of scale being important for achieving good generalization. So not just flexibility, the simplicity of bias that comes about through scale. I think another misconception is this idea that we should change our model depending on how many data points we happen to have available. And this might even be the most controversial 1. The reason I don't think we should is because we should always honestly represent our beliefs. And our beliefs about the process that generated our data typically shouldn't change depending on how many data points we happen to have access to. And you can actually demonstrate that these principles work in practice. So you can have models that will be very good when you have a small number of data points, and also very good when you have a very large number of data points. And so this relates to not necessarily needing to have hard constraints, but instead combining expressiveness with Occam Tracer. Yeah. And you know, you're you're only slowly starting to convince me to give up this 10,000, you know, degree polynomial is bad. And and and I'm only starting to change because you value simplicity, bias towards simplicity as much as I do. And the real key for me was understanding that somehow scale has a bias towards simplicity. And I don't know why or where it comes from,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.