Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“These are continuous functions because most of the time in a neural network, like, if you if you give it a test sample, which is outside of the training support, it you're in no man's land.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. That's a that I think maybe that's a bit unknown for me. I'm not sure because I think there's there's trade offs there. If if there if it's reduced again, it's this Goldilocks thing. It's like if you reduce it too much, then it ends up sort of not being able to do the things that you need it to do. And if there's too many parameters, it's able to do the things you need it to do plus it does a bunch of stuff like you definitely don't want it doing, you know. And so there's just this weird, you know, this weird balance in there. And and yeah. So I think that's a very difficult question and and very hard engineering problem. But this paper is pointing in in directions that we need to go and actually 1 really great thing about this paper which you mentioned earlier is how visual it is. So like if you look for example at the at the the intermediate layers for the evolved, you know, compositional neural network versus the SGD trained 1. Okay? Like, I mean, come on, folks. Like, I I don't care how much of an SGD fanboy you are out there. Like, just look at the images and come and tell me with a straight face, okay, that there's any merit whatsoever to the entangled, convolved, contorted, convoluted nonsense that, like, you end up with, like, with actually, like, seriously, like, I don't even know if we have a conversation if you think there's, like, value in there. But if you because if you look at the evolved ones, they're like a stencil that you would could imagine an artist would have. Like, well, here's an oval thing and some, you know, gradient lines and and this sort of thing. And the point is it's not just about appeal to humans. Okay? It results extremely powerful transformation capabilities like these parameter sweeps where there's a parameter that controls the mouth opening and closing or the mouth being wider or winking or the apple's, you know, leaf literally rotating as if it was on a 3 d axis. Okay? So these are neural networks that were trained solely on 2 d images and they have evolved a concept of, 3 d rotation. Like, if you're not impressed by that and kinda blown away by how cool that is and if you can't appreciate that that's relevant to the real world and agents operating in the real world with an effective world model, like, I don't know. I can't really pull you out of that hole. Yeah. And I wrote some software to visualize that so we can show that on the screen. But this actually comes back to the famous episode 61 of MLST when we had Randall Bellestriero and and Jan Lecun. It was called interpolation extrapolation something else. I can't remember now. But yeah, we we showed this neural network visualizer, know, was the it was the TensorFlow playground, think it was called. And it it would show you like the the evolution of of training, you know, for all of these different types of neural networks with different act activation functions and learning rates. And 1 of the data sets was a spiral manifold. So it'll show you a spiral and you could kind of see it and it was it was trying to fit the the the thing over time. And 1 of the reasons why it can only interpolate when inside the training support, which basically means roughly speaking inside the, you know, there's this high high dimensional space of data and it can roughly, you know, go inside the space. It's because the activation functions are are re reload. Right? They're they're they're piecewise linear functions, which means they just chop up the space. Now, what's interesting about the the evolved CPPN, the compositional pattern producing networks, is that the functions can actually be trigonometric functions. These are continuous functions because most of the time in a neural network, like, if you if you give it a test sample, which is outside of the training support, it you're in no man's land. It gives you nothing. Whereas with these CPPNs, if I understand correctly, they actually do, extrapolate outside of the training support because these trigonometric functions, they just they just keep going. Right? The problem isn't just the activation functions. It's the training methodology. And and what would happen, for example, let's say we just keep rel use. I think if you still found a way to do this kind of evolutionary building up from simpler, you know, mappings, what you would end up doing is there would be there would be kind of an a a simple higher order layer that say took the spiral and chopped it up into 4 quadrants, like nice nice quadrants. And then 1 that did sort of like some 45 degree angle chopping and whatever. So it would still end up being a piecewise linear division of the space. But rather than if you go and do this over at the TensorFlow, you know, site, it's all messed up weird, you know, lines, funky angles, and whatever. Instead of that, you would end up with the type of chopping up that you and I might do. Right? Like quads and then hex hexagons and like whatever else. So the point is whatever activation functions you have, there's a better way to train them.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.