Evidence receipt / recommendation
Published · transcript-backedAndrew Gordon Wilson: recommendation
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“Where like you have outliers but they're different each time. So just training on the outliers isn't gonna be useful because they're gonna be new outliers that look very different from those outliers.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…iving them access to to a calculator. So why don't we just do that? Or like, you know, is there some sort of greater value in having to sort of learn a representation that can do some of these things? Because maybe even though we'll always wanna use a calculator if we're multiplying or adding numbers, being able to do something like that reasonably well might transfer into other settings that we don't really anticipate. And I'm probably more in this category, like I just think intellectually we should try to figure out how to do this without tools. And so this work that we're doing on transformers for matrix operations, I think is quite analogous. So you could argue that addition and multiplication are just fundamental primitives for trying to build intelligent systems. We're not as good as a calculator is, but it might be important for us to have some ability at being able to do these things. The same could be said of matrix operations. This is really the backbone of all sorts of different learning algorithms like Gaussian processes, which we discussed a little, involve solving linear systems with a covariance matrix, computing log determinants, etcetera. Deep generative models like normalizing flows involve log determinants, dimensionality reduction like PCA, etcetera involve other matrix operations. And so they're just really ubiquitous as kind of a primitive for trying to build learning algorithms and intelligent systems. And so I think if transformers are ultimately going to become some sort of general intelligence, then they ought to be able to to be confident at these types of operations. And so we were kind of representing then matrices as sequences of numbers and then having the outputs be things like the maximum eigenvalue solution to a linear system or whatever other operation we were considering, the spectrum of eigenvalues. And we found interestingly that when you train this approach on Gaussian random matrices, you just sample every entry of the matrix from a standard normal distribution. It will do fairly well at in distribution matrix operations, so other matrices sample from that distribution that it hasn't seen before, but extraordinarily poorly, even if you go slightly outside of that distribution. So given an identity matrix, all ones in the diagonal, 0 everywhere else, that'll have pretty low density, has support under this Gaussian distribution or matrices, but pretty low density under it, hasn't seen something very much like that. It will just completely fail. It won't do anything reasonable. And so we considered a number of interventions, like looping, sort of adaptive test time computation, enriching the training distribution very significantly, so having sort of this EinSum space of structured matrices that we were sampling from, so all sorts of different matrix structures, Toplitz, Kroniker, etcetera, block diagonal, low rank, and so a variety of other interventions. And interestingly we found once we had done that, this approach actually was able to generalize even to matrices that were out of distribution for this fairly rich sort of training set that we've created. interestingly we found once we had done that, this approach actually was able to generalize even to matrices that were out of distribution for this fairly rich sort of training set that we've created. And so it seemed to move more towards learning an algorithm rather than just doing statistical interpolation on the training data. And so this is I'm an optimist by nature, so I I am very happy about the fact that we can use data to move towards things like algorithm discovery. But now I remember why I started talking about this. So I think there are instances where it's going to be hard. So autonomous driving is is an example Right. Where like you have outliers but they're different each time. So just training on the outliers isn't gonna be useful because they're gonna be new outliers that look very different from those outliers. And I just have the intuition that more data alone is not really the answer to building robust autonomous driving systems. Yeah. I just wanna say as far as why not just give them tools, I mean, my answer folks is because those have to be programmed and built by people. And the whole point here is to allow machines to do their own programming. Right? Machine learning. And if they can't if they can't even learn to do multiplication reliably and to generalize from decimal multiplication to binary or hexadecimal and 9 digits to 36 digits. You know? What hope do we have that they're gonna discover relativity or non Newtonian mechanics or any other kind of frontier frontier things? Right? I mean, isn't that part of the goal here?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.