Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“Right? And we are different from that because, you know, the the very basis of how we think is correlated to how the world works.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Right. Well, I mean, in the spirit of his title, you know, questioning representation optimism, I I think, you know, Kenneth has engineering optimism. And and I like, to a degree, I share that optimism. So, I mean, you know, because because you and I both have spent enormous amounts of time in our Discord server arguing about things like this. So I'm well intimately at a gut level aware of just the absolutely cosmically insane amount of computation that goes into powering the the algorithm of evolution with a capital e. You know, it's happening on the earth in many forms, not just life, but memetics and information and, you whatever else. Right? Higher order phenomenon. So I'm just saying there are probably very clever hacks. At least I hope, maybe Kenneth hopes that there are clever hacks that will get us somewhat towards a better method, you know. Akarsh had a very good take on on kind of, you know, some of these evolutionary aspects in particular, you know, it it not being a an optimization but a constraint. And so it's good to see, you know, new new researchers. Right? Like kind of taking up taking up that mantle. And I think a big part of what the point of this paper is and their point is we need to think more about these processes because I think there are, you know, like you brought up a lot of interesting, you know, sort of takes just now and intuitions on how POET and these algorithms might work. And it's I think the building up is is essential. So I think what we need to do like, imagine some future training neural networks will will look like this. Like you'll start with a neural network that has a 100 parameters and you'll train it on some some subset of the data, some simple subset. And then the next phase, you'll train 1 that has 1,000 parameters And in some way, it's kind of expanded from that lower level network. Like, we'll have algorithms that will identify sub networks in there that do something in particular. Let's say these 12 neurons. And what we'll do is, like, we'll now make those 120 neurons and then train them and and keep them sort of isolated as a group. Right? So that they grow that abstraction, like, sort of further. And then it'll be 10,000, a 100,000. Like, you'll kinda keep scaling it up. Right? And then what you'll end up with is almost like a neural network that has a very specific kind of not sparsity, but modularity. Like a very specific kind of modularity because it was grown from like seeds, you know, up rather than starting at the massive scale and allowing it to do all this like really obnoxious and and counterproductive entanglement. Yeah. And this is the I mean, I think we're getting to putting the pin in in the center of the dartboard here. So he said the the problem is that neural networks have too many degrees of freedom. They're like a pile of sand. Right? And we are different from that because, you know, the the very basis of how we think is correlated to how the world works. So there's something about building up, you know, you said building up not not tearing down. And and he said it's not about what you know, it's about how you got there and how evolvable the knowledge is. So where you can go with that knowledge. And it's not necessarily that stochastic gradient descent is a bad thing in of it itself. It's simply that it's statistically intractable to use anything other than stochastic gradient descent. Otherwise, it simply doesn't work. So he's not necessarily saying, I mean, maybe he is, but he's saying that these evolutionary algorithms are really, really good when you're building up, when you have a very, very sparse search space. And and then the other thing which is very important which we didn't get to is this whole building up thing like the neat algorithm and you can bring that in. Is this idea that we need to have a form of training or learning about the world which monotonically increases information and complexity. So we're we're adding this thing and we're adding this thing. And every time we compose things together, we're looking at the evolvability. So we're not just saying, can you do the thing now? We're saying, is it actually likely to be able to deal with future things that I might encounter? When children are born, they still have 2 legs, basically. So you shouldn't just think of it as a random, you know, crossover operation. It's an operation that respects the the provenance, the topology, the structure. Right? So there are, you know, in practice, there are crossovers and mutations within certain topological frames. But some things are held constant because they need to be held constant. But on your comment about degrees of freedom, I think Kenneth was saying that we actually want to reduce degrees of freedom. So he's not kind of saying, oh, we just need degrees of freedom for stochastic gradient descent, but degrees of freedom are okay. I think he was saying that there's actually a magical Goldilocks zone when it comes to degrees of freedom. So we want the degrees of freedom to be at least a representation of how the world works. But necessarily, there should actually be more flexibility than how the world works. But if there's too much flexibility, you get a weird kind of mode collapse and you don't see the emergence of evolution. Degrees of freedom, the way I think about is, like, it's like a memory with an algorithm. You want to use as much as necessary and no more. And it's really hard to it's really hard to figure that out.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.