Evidence receipt / uncertainty
Published · transcript-backedTim Scarfe: uncertainty
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“Duggar. But some people say, I don't know, like polysemanticity or grokking or scale and, you know, that that it just it just appears like the neural network isn't grokking it.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Sounds good. That's a good Copyright, Keith. Duggar. But some people say, I don't know, like polysemanticity or grokking or scale and, you know, that that it just it just appears like the neural network isn't grokking it. But actually, it is. And maybe if you just trained it a little bit differently. Part of what's important is just to point out that we've identified a really interesting phenomenon, but we don't have answers to all these questions yet. And sort of part of what we're doing is saying, let's go find out the answers to these things. It's possible that some of them help. It's possible some of them don't help. I think 1 thing just to think about, just to observe upfront is that the difference that you see visually is so dramatic and stark. Anybody who looks at the figure of the 2 different versions of the skull, it's hard to believe, although we don't have the evidence yet, that there's something that's going to fix this that's just real simple, like a grok. I imagine grokking does something good, but it's hard to believe it ends up making something that looks like the CPPN version of the skull from Pick Breeder. That's hard to believe. But nevertheless, it's possible that things like rocking and all kinds of other things that we do, mixture of experts, maybe convolution helps in some ways, although that's not in the LLMs, but it could be in other networks helping. They're all just questions that are not answered. But I think 1 thing to think about with respect to things like grokking is just, even if you're trying to rationalize and save your faith in how these things work right now, and think, Oh, well, it's okay, grokking will fix this. Just think about this, like wouldn't it be nice if you didn't have to do grokking? What grokking basically means is that you get this absolutely horrible entangled mess and you clean it up later, like once you start figuring things out, start deleting all the excess and redundancy, all the fracture gets fixed. Wouldn't it be nice if you don't have to do that and it's just good in the first place? What we've seen because of the PickBreeder CPPN is that that actually is possible. Without this example, without this weird example from PickBreeder, almost no 1 would believe or buy that that's even possible. It's like, well, I mean, of course I'd love miracles to happen every day, but it's just not gonna happen. But the thing is that it did happen and it happens consistently. So it raises this question is that even if grokking is helping, is there a better way or maybe a complementary way? Because maybe both factor into why human brains get to where they get to. So these are unanswered questions. This is not to say that grokking doesn't help. We need to check and find out the degree to which it helps. But 1 other kind of interesting thing to think about is that if you think about the efficiency of training, how expensive it is, like we're building, spending billions, hundreds of billions of dollars, creating these giant data centers to do giant kinds of training runs. Like, if it's true that there's an incredibly more efficient way of doing things. So for example, it could be that, yes, SGD with a brute force will do basically almost anything at least in terms of being able to get the training data.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.