Evidence receipt / belief
Published · transcript-backedAndrew Gordon Wilson: belief
19 Sept 2025 Machine Learning Street Talk Deep Learning is Not So Mysterious or Different - Prof. Andrew Gordon Wilson (NYU)
“Absolutely. And I think being able to do certain things while we're finding might surprisingly relate to doing other things very well.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon Wilson
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 19 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. I just wanna say as far as why not just give them tools, I mean, my answer folks is because those have to be programmed and built by people. And the whole point here is to allow machines to do their own programming. Right? Machine learning. And if they can't if they can't even learn to do multiplication reliably and to generalize from decimal multiplication to binary or hexadecimal and 9 digits to 36 digits. You know? What hope do we have that they're gonna discover relativity or non Newtonian mechanics or any other kind of frontier frontier things? Right? I mean, isn't that part of the goal here? Absolutely. And I think being able to do certain things while we're finding might surprisingly relate to doing other things very well. And so I think we have seen this to a large extent with LLM. So we had a paper where we just took a text pretrained LLM off the shelf and applied it to time series forecasting. So we did everything naively. We weren't even really intending for this to be like a proper project or a paper. We were just curious, you know, if you just give GPT like a sequence of numbers naively encoded to strings and have it extrapolate the next sequence of string tokens, how would it compare to purpose built time series forecasting procedures? And it just worked way way better than we thought it could possibly work. It didn't even really make sense. And so we did sort of turn this into a proper project. We did a little bit of work on trying to improve the tokenization and think about uncertainty representation, etcetera. But most of that paper, it's called Large Language Models or 0 Shot Time Series Forecasters, was focused around trying to understand how this is even possible. And in the end, it it did start to feel a bit more like maybe it's not just that you can do this, maybe you should do it in some instances. Like they did quite well on on a variety of benchmarks. And to me, this suggests with like a proper dedicated research effort on LLMs for time series, you know, we could see these systems actually working a lot better than the purpose built models. And so what that shows is being able to predict the next words in sentences can actually transfer to being able to do other things like series prediction really well. And we had another paper on LMs for materials generation, which was kind of similar. So we took a text pretrained LLM off the shelf, like a LAMA 2 model at the time. And we fine tuned it, in this case, on atomistic data represented as text, so locations of atoms and energies and things like this. And the resulting system was able to generate inorganic crystals with favorable properties better than these purpose built approaches, and even foundation model approaches that had been trained on that domain specific data. And so 1 of the takeaways from that project was like the text based pre training was an indispensable component in being able to achieve good results on materials generation. And we tried to understand also in that paper why that was the case. We were collaborating with some some chemists at Fair in California, and they were just very curious about LMs and foundation models. So they're willing to sort of humor us a bit and help us sort of see what we could do. But they were just like skeptical throughout the project until we we saw the results. And it's like, well, you can't deny the results are great. You know, why is this happening? And part of it was that in being able to predict the next tokens in strings, you are learning principles of induction like Occam's razor. t deny the results are great. You know, why is this happening? And part of it was that in being able to predict the next tokens in strings, you are learning principles of induction like Occam's razor. And so how do those sort of manifest themselves in in context learning? So what it means is if you are and in fine tuning. So what it means is these models are gonna be predisposed to discovering compressible representations. And that means, for example, salient symmetry. So this was a problem where there was a rotation invariance. And these models actually were very quick to learn these kinds of invariances because of the text based pre training. And it sort of instilled this principle. And so I think this is also an example of how compression can be a broadly applicable principle for induction. It's after the right level of abstraction that you can start to see more relatively universal behavior. So for instance, there are some theories on the success of foundation models that suggest that different problems are just different projections of some underlying reality. So like platonic representation hypothesis is an example of this. Mhmm. And so you can represent an image with pixels or with words and you know, train the respective models on those different modalities and they learn similar representations. I think that can be true in some instances, but I also think different problems are often truly quite different from each other in terms of their low level structure. So like if we have molecules, there's rotation and variance, maybe some other problems, some image recognition problem, like character recognition, we might have translation invariance. Doesn't matter if the 2 is on the left of the screen or the right of the screen, the label is still a 2. These are very different low level feature representations. So like the architectures that you would typically use for each of those modalities and to respect those different types of invariances would look very different. But what they have in common with each other is they're both ways to compress the respective problems that they're being applied to. So if you have a model that has this compression bias, then it can discover those salient symmetries. And we had this really surprising finding that vision transformers actually can be more translation equivariant than convolutional neural nets after training, which sounds impossible because conv nets by design are are translation equivariant. But they're not exactly translation equivariant because of aliasing artifacts and edge effects and things like this. And so this other model, this transformer with no explicit constraint whatsoever, just a soft bias that manifests itself increasingly at scale, is able to discover a solution that has lower equivariance error than the convolutional neural net, which is just absolutely remarkable. So to come kind of full circle to your question, I think that we're discovering more and more that being able to solve certain problems really well will translate in perhaps unexpected ways to being able to solve other problems well.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.