High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Taco Cohen: evaluation

22 Dec 2025 Machine Learning Street Talk Making deep learning perform real algorithms with Category Theory (Andrew Dudzik, Petar Velichkovich, Taco Cohen, Bruno Gavranović, Paul Lessard)

“In my PhD, I worked a lot on building, knowledge about symmetries into neural networks. And I think, for for many problems, you know, knowledge about symmetries is something that, first of all, gives you a lot of bang for the buck.”

— Taco Cohen

Source trail

Everything needed to verify it.

Speaker
Taco Cohen
Attribution
Verified speaker
Claim type
evaluation
Recorded
22 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…formers, at their heart, are permutation equivariant models. Once you've put token embeddings, position embeddings into tokens, you can permute them all you want. You'll get exactly the same response. If you wanted to learn that kind of symmetry with a simple MLP of tokens, that would have taken you exponentially more data than the trillions that we currently use to train these models, so likely a data you wouldn't be able to find. We looked at geometric deep learning from a a group symmetry point of view, which is a very nice way to describe spatial regularities and spatial symmetries, but it's not necessarily the best way to talk about, say, invariance of generic computation, which you would find in algorithms. Right? It's like I have input that satisfies certain preconditions. I want to say something about once I push it through this function, it should satisfy certain post conditions. This is not the kind of thing we can very easily express using the language of group theory. However, it is something that perhaps we could express more nicely using the language of category theory. I do think that some very high level priors, are probably a good idea and perhaps even necessary. In my PhD, I worked a lot on building, knowledge about symmetries into neural networks. And I think, for for many problems, you know, knowledge about symmetries is something that, first of all, gives you a lot of bang for the buck. You you we know from from physics already and now from empirical results in in machine learning that building these things into new neural networks or putting a constraint on your physical theory based on symmetries gives you a lot of information or it really restricts the space of of hypotheses. And at the same time, it doesn't bias your model if indeed your problem has this symmetry. So I think that that that kind of that's the kind of thing we should be be looking for. This kind of very high level abstract prior, not trying to encode, you know, go coming back to the example I gave just now, the fact that light switches make lights go on, that we can figure out from data, from reading text on the Internet at scale, from trial and error, learning in an interactive environment. Perhaps the fact that there is space, like 3 d space, and that your 2 d images are a projection of that, maybe that that is a useful prior. Category theory is very much in the eye of the beholder. I think in the first instance for me, theory categories are a very mundane thing from pure mathematics where I come from. And category theory means, you know, when you study categories for their own sake. But everybody uses categories. The question is what exactly are they? And I really come from algebra, and a lot of my motivation comes from studying algebra. And 1 way you can think about categories is algebra with colors. So, you know, we can imagine sort of typical algebra. So for example, let's say we're multiplying square matrices. I can sort of think of each square matrix as a little magnet, and I just sort of hook them up together, and they just stick and you kind of, you know, you get a bigger and bigger magnet and and everything makes sense. But now suppose I had special magnets that had colors on each side and I could only connect them if the colors were the same. And that sounds a bit weird but it's exactly what happens with non square matrices. When we, multiply 2 matrices, we we have to follow a rule that they are not allowed to be multiplied unless the numbers match up. If I have an m by n matrix and I wanna multiply on the left with an l by m matrix, I can do that because the m's are the same. But otherwise, I I can't. There's a kind of color violation. So the point is a situation where we want to be able to compose things to to hook them up together, but we can't always do it. That's basically what categories are designed to cover, and I think the the matrix example illustrates they're not so mysterious. It's just when you want to be talking about, for example, many different sized vector spaces at once, as you often do in neural networks because you have sort of hybrid shapes of with, you know, dimensions of of different sizes and so on, you end up you end up wanting something where you take this sort of partial compositionality into account.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence