Evidence receipt / belief
Published · transcript-backedKeith Duggar: belief
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“I think if you still found a way to do this kind of evolutionary building up from simpler, you know, mappings, what you would end up doing is there would be there would be kind of an a a simple higher order layer that say took the spiral and chopped it up into 4 quadrants, like nice nice quadrants.”
Source trail
Everything needed to verify it.
- Speaker
- Keith Duggar
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. And I wrote some software to visualize that so we can show that on the screen. But this actually comes back to the famous episode 61 of MLST when we had Randall Bellestriero and and Jan Lecun. It was called interpolation extrapolation something else. I can't remember now. But yeah, we we showed this neural network visualizer, know, was the it was the TensorFlow playground, think it was called. And it it would show you like the the evolution of of training, you know, for all of these different types of neural networks with different act activation functions and learning rates. And 1 of the data sets was a spiral manifold. So it'll show you a spiral and you could kind of see it and it was it was trying to fit the the the thing over time. And 1 of the reasons why it can only interpolate when inside the training support, which basically means roughly speaking inside the, you know, there's this high high dimensional space of data and it can roughly, you know, go inside the space. It's because the activation functions are are re reload. Right? They're they're they're piecewise linear functions, which means they just chop up the space. Now, what's interesting about the the evolved CPPN, the compositional pattern producing networks, is that the functions can actually be trigonometric functions. These are continuous functions because most of the time in a neural network, like, if you if you give it a test sample, which is outside of the training support, it you're in no man's land. It gives you nothing. Whereas with these CPPNs, if I understand correctly, they actually do, extrapolate outside of the training support because these trigonometric functions, they just they just keep going. Right? The problem isn't just the activation functions. It's the training methodology. And and what would happen, for example, let's say we just keep rel use. I think if you still found a way to do this kind of evolutionary building up from simpler, you know, mappings, what you would end up doing is there would be there would be kind of an a a simple higher order layer that say took the spiral and chopped it up into 4 quadrants, like nice nice quadrants. And then 1 that did sort of like some 45 degree angle chopping and whatever. So it would still end up being a piecewise linear division of the space. But rather than if you go and do this over at the TensorFlow, you know, site, it's all messed up weird, you know, lines, funky angles, and whatever. Instead of that, you would end up with the type of chopping up that you and I might do. Right? Like quads and then hex hexagons and like whatever else. So the point is whatever activation functions you have, there's a better way to train them. Yeah. I mean, there's there's a few points here because I guess like what I was, where I was going with this before is, you take y equals x squared, and the reason why we think of it as robust is for any value of y, it kind of does something it does something reasonable. Right? It's not it's not just memorizing little little regions in in the training support. We often speak about the the limitations of neural networks in a variety of different ways. So there's there's complexity limitation. There's a computational limitation. There's a training limitation. But it's quite interesting to play with this and pull it apart a little bit. So yeah, as you say, the SGD thing, that is actually a serious limitation because there are neural networks that can do a hell of a lot more if they are trained with SGD. Mean, Schmidhuber gave an example of, you know, the that it's possible it's possible to have an RNN which is Turing complete under certain conditions, but it's not trainable with with SGD. And there's also the other matter that when you start using these like weird and wonderful activation functions, you can do things outside of the training support. And then there's the other matter that when you start doing non SGD training, you can actually build hybrid systems as part of the training process. So imagine if you if the way we recognize faces, we we do some kind of composition, splitting, iteration, recursion. What you could just train that entire system with another form of training process. So it rather opens up the whole game.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.