Evidence receipt / evaluation
Published · transcript-backedKenneth Stanley: evaluation
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“We also show new examples. But you you know, the word reason I use the word indirect is because unlike these pick breeder images, we can't just go in and look at a neuron and know what it does explicitly because that's what's so nice about pick breeder images because they're 2 d.”
Source trail
Everything needed to verify it.
- Speaker
- Kenneth Stanley
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yep. Yeah. And I mean and, also, I think there's a vast underappreciation of of the fact that evolution occurs in many scales and many systems. Like, it occurs in psychology. It occurs in in means. It occurs in the information space. Like, it it's a very general, almost meta algorithm. Right? And and I would encourage people to look into whatever domain they're in, whether there's any kind of evolutionary perspective on it because it can be very fascinating. Like, a lot of human behaviors that people dismiss as irrational are actually not irrational if you consider it from the perspective of an algorithm operating on a ecosystem of species or individuals or or things like that. So it's just fascinating. The pick breeder is like that. You know? I mean, the pick breeder has a lot of evolutionary properties, and here it is creating a neural representation that has these fascinating properties internally, and that can't be dismissed. I mean, the the evolutionary properties of what the human beings are doing in pick breeder is obviously part of the explanation. But it's not that we want to have to depend on humans guiding the search process. It's whatever they're doing that's somehow analogous to what evolution did, which allows it then to discover these incredible internal representations. We wanna abstract that. So we wanna find, like, yeah, the the correct level of abstraction to allow this to happen in in autonomous fashion so that the training regimens of these large models leads to a similarly amazing underlying representation. And I think that the question that looms is like how far is the representation currently from this ideal? Because it's it's not totally known. Like, we show in the paper a number of pieces of evidence, indirect evidence that there is fracture inside of fracture and entanglement inside of today's large models. We we we go through literature. We also show new examples. But you you know, the word reason I use the word indirect is because unlike these pick breeder images, we can't just go in and look at a neuron and know what it does explicitly because that's what's so nice about pick breeder images because they're 2 d. You can just look and see exactly what's represented. It's more implicit in like a real giant multi, you know, multibillion connection network. And so the easiest thing to do is just look at indirect evidence of of fracture. But there's this overriding lesson here that even when you're getting answers right, you could be messed up under the hood, which then, you know, further confounds our ability to analyze what these models really know. Because, like, even if they got everything right, they could be just like the the the network that that outputs the skull, which is basically almost perfect, or it is perfect. And so it it could be perfect, every every question we ask is perfect, and it could still be an absolute mess underneath the hood, which means that this indirect evidence, you know, could diminish over time, like, as the models get so called better, but it might not mean that they're fixing the underlying representation to be less fractured and entangled. I think it's super important that you you name this concept impostor intelligence in the paper. So I think it's important that we that we point that out. Like, the idea that if sure. If you have enough fractured, entangled representations, you can perfectly reproduce an output, you know, that looks identical to the 1 that you get from unified, oh, shoot. What are they called again? Unified Factored. Factored representations. Right? Can look identical on the surface. The output can can look identical, and that's impostor intelligence. Correct?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.