Evidence receipt / belief
Published · transcript-backedTim Scarfe: belief
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“I mean, there's there's a few points here because I guess like what I was, where I was going with this before is, you take y equals x squared, and the reason why we think of it as robust is for any value of y, it kind of does something it does something reasonable.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…training methodology. And and what would happen, for example, let's say we just keep rel use. I think if you still found a way to do this kind of evolutionary building up from simpler, you know, mappings, what you would end up doing is there would be there would be kind of an a a simple higher order layer that say took the spiral and chopped it up into 4 quadrants, like nice nice quadrants. And then 1 that did sort of like some 45 degree angle chopping and whatever. So it would still end up being a piecewise linear division of the space. But rather than if you go and do this over at the TensorFlow, you know, site, it's all messed up weird, you know, lines, funky angles, and whatever. Instead of that, you would end up with the type of chopping up that you and I might do. Right? Like quads and then hex hexagons and like whatever else. So the point is whatever activation functions you have, there's a better way to train them. Yeah. I mean, there's there's a few points here because I guess like what I was, where I was going with this before is, you take y equals x squared, and the reason why we think of it as robust is for any value of y, it kind of does something it does something reasonable. Right? It's not it's not just memorizing little little regions in in the training support. We often speak about the the limitations of neural networks in a variety of different ways. So there's there's complexity limitation. There's a computational limitation. There's a training limitation. But it's quite interesting to play with this and pull it apart a little bit. So yeah, as you say, the SGD thing, that is actually a serious limitation because there are neural networks that can do a hell of a lot more if they are trained with SGD. Mean, Schmidhuber gave an example of, you know, the that it's possible it's possible to have an RNN which is Turing complete under certain conditions, but it's not trainable with with SGD. And there's also the other matter that when you start using these like weird and wonderful activation functions, you can do things outside of the training support. And then there's the other matter that when you start doing non SGD training, you can actually build hybrid systems as part of the training process. So imagine if you if the way we recognize faces, we we do some kind of composition, splitting, iteration, recursion. What you could just train that entire system with another form of training process. So it rather opens up the whole game. Right. Yeah. A 100%. I mean, like and you can handcraft neural networks. Like, for example, you know, there was that was that a DeepMind paper. Right? The the 1 where they handcrafted sort of a universal Turing machine in a in an LLM form, but you also have to add to it, you know, expandable memory. So they had to do extended generalized autoregression rather than just, you know, autoregression. So that little technical detail aside. Like you're saying, you can handcraft things, you know, you you can construct them, but they're not trainable. And so it's the core thing we need to improve is the training methods. And I think nobody knows what they look like really, like at the moment. Like there are some hints. Right? Like POET, maybe what you're talking about, these kind of hybrid systems where things are trained in different ways or the curriculum learning or some type of building up, you know, from seeds that are somehow kept slightly isolated. I don't know. Like, I I I we don't know because everybody's just stuck on SGD SGD scale scale scale is enough. SGD attention's enough. It's like until the industry kind of fails enough to achieve anything real like with those methods, don't think they're I don't think sufficient resources are gonna go into the alternatives. Yeah. They're getting close. They're getting close cause even the doom even the doomers are having conversations about like why aren't we all dead yet? Like we thought, you know, once GPT 4.5 came out or 4 point whatever, it's like I thought we were doomed then and like what's what's going wrong? So they're kind of reevaluating their, you know, apex of rationality over a less wrong to figure out what went wrong. You know, why aren't why aren't we dead?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.