Evidence receipt / evaluation
Published · transcript-backedKenneth Stanley: evaluation
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“It's just emergent from how SGD climbs these gradients. But the thing that I think makes this really intriguing, the reason that the paper is because beyond just that, is that it gives you something that otherwise could never exist, which is a counterexample, that there actually do exist networks that don't have that issue.”
Source trail
Everything needed to verify it.
- Speaker
- Kenneth Stanley
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So it goes back to these observations from Pickbreeder, which is this picture breeding website where neural networks, which were a special kind of neural network called compositional pattern producing networks, were bred or evolved by people to produce images. And this is a really unusual and exotic form of training that's very different from modern deep learning. It's completely different from it, which is 1 of the important points about it. And so when people were able to actually breed interesting looking pictures, there's a lot of lessons we learned from that. Actually, the novelty search algorithm and things like that came from those lessons. But there was this 1 dangling lesson that, like, just hasn't gotten out, and that's what this paper does. And that that lesson was that the representations, the underlying representations of these images, which are basically represented encoded by these neural networks are absolutely incredible amazing. And there was like no good explanation for how they could be as good as they are. And we noticed this very early on that like they have unbelievable modular decomposition, which means that it's almost like it was engineered by a person. You know, it's like there's a there's there's a network that that generates the image of a skull, and it the network has decomposed it such that, like, there's a component of the network that's responsible for the mouth. They can do things with the mouth, like open and close the mouth, or there's another dimension that can make the mouth smile. And things like that which are not it's what's really amazing. It's not data driven. Like, we're talking about dozens of iterations in terms of, like, the amount of search that's involved to get a representation like this. And so I've always there's we've had many discussions over the years, like, why why is the representation so amazing in these pick breeder images? Why does it have this unbelievable modular decomposition? And and and extremely efficient and compact. And so 1 thing that like at 1 point I did with Joel Lehman was we were kind of just playing around to see, like, well, what if we tried to get SGD to produce the same images? Like, what what would it be different internally? The the image the way that it represents these images. And it was just dramatically different. But then we did nothing with that. Like, we just kinda, like, left it. We both Joel and I knew this thing, but we didn't really do anything with that information. But I've been just mulling it for years that, like, there's something there that's really important That, like, with this weird kind of open ended search process, you get these incredible representations. And then with conventional, like, objective driven SGD, which is, the the backbone of all of machine learning right now, you get a completely different kind of garbage representation, just total spaghetti. hen with conventional, like, objective driven SGD, which is, the the backbone of all of machine learning right now, you get a completely different kind of garbage representation, just total spaghetti. And we we we used we came up with some terminology that we put in the paper to to more, like, clearly articulate what these differences are. But basically, you're talking about just amazing versus garbage. And so like the the question that the paper addresses is what does this mean? Which is something I think that has endless repercussions and potential implications. Like the fact that we're basing the entire field on something that produces this complete garbage under the hood. And does this mean anything? Now I think if if you just saw this by the way, the fact that it's images makes it easier to see, like the the I mean, that the representation is garbage. Like the fact that it's a network that produces 1 image. Because what that means is that we can visualize every single neuron, what subcomponent it's actually computing. It's very easy because it's just a single image. So it makes it very easy to see all at once how it actually computes the image, so we can see the internal representations. And if I just showed you that for SGD, like if I said, here's a skull, I made this network, reproduce this skull, and the internal representation is crap. I think it would be not like hugely resonant in the field, because it'd be like, oh, it's not super surprising that it's kinda hard to understand this like really complicated entangled representation. It's just emergent from how SGD climbs these gradients. But the thing that I think makes this really intriguing, the reason that the paper is because beyond just that, is that it gives you something that otherwise could never exist, which is a counterexample, that there actually do exist networks that don't have that issue. You would think that that's just intrinsic to neural representation, that somehow they just look like kind of entangled messes, and that's just the way life is. But clearly, it's not how life has to be. And so because we have now actual concrete examples from Pic Breeder that show beautiful internal representation. And so I think that the the paper's trying to get at what does that mean? Like, how how should we interpret this? Should we perhaps aim for algorithms that actually do achieve these kind of, you know, really amazing kinds of modular decompositions? Or should we or or should we just say, well, no, that's that's just not important. Let's just forget it and dismiss the fact that we know they exist. I think that would be that seems to be a pretty premature move to say that. So so there's something here, think, deep deep lesson for us to learn in the field about internal representation, which just is totally off the radar. It's way way off the radar of anything that's being discussed. You know, mechanistic interpretability is not on the same page with what we're showing here. And so, of course, it it can play a role in in further, like, interpreting this result. But this is a new thing to look at and for us to contemplate and try to understand what the implications are. Don't we just need to have humans in the loop? I mean, wouldn't it be great if we could just have these evolving topologies, right, without humans? Wouldn't it be great if the systems could themselves learn the categories, the abstractions? And is it just a matter of missing data? Or is there just a fundamental gap in capability?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.