Evidence receipt / belief
Published · transcript-backedKenneth Stanley: belief
6 Jul 2025 Machine Learning Street Talk The Fractured Entangled Representation Hypothesis (Kenneth Stanley, Akarsh Kumar)
“Like, what are you objecting to? But the but the point is that it can still be an impostor because it's like what we care about here is not just that it's going to get answers right, like get good test scores, like seem to be plausibly human when you talk about things that are in distribution.”
Source trail
Everything needed to verify it.
- Speaker
- Kenneth Stanley
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 6 Jul 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I think it's super important that you you name this concept impostor intelligence in the paper. So I think it's important that we that we point that out. Like, the idea that if sure. If you have enough fractured, entangled representations, you can perfectly reproduce an output, you know, that looks identical to the 1 that you get from unified, oh, shoot. What are they called again? Unified Factored. Factored representations. Right? Can look identical on the surface. The output can can look identical, and that's impostor intelligence. Correct? Yeah. Yeah. Yeah. It it just seemed like a like a good another good metaphor is to think of it as an impostor. It's like this you know, the the representation of the skull is just somehow a farce. Like, it's not what you expect. Like, if you just look at the output, it's great. It looks exactly like a skull. But underneath the hood, it's it's not capturing any of the underlying components or the regularities. And so in some sense, it's not really a skull. It's an impostor underneath the hood. It's very misleading. And so if you think of the skull again as a metaphor for all of human knowledge, because that's what an LLM is trying to capture, it's not just a single image, it's like an image of all of human knowledge. For any input, it should output something that's convincingly human, then like it it could be just similarly an impostor. It could be the same. Underneath the hood, everything could be organized wrong, not the way you expect. It's like a giant charade. And again, this is very confusing and counterintuitive, like, for for for people because, like, you know, people are naturally, including even me, I would react like, but should I really care? Like like, when you say it's an imposture, but it's getting everything right in its human level. Like, what are you objecting to? But the but the point is that it can still be an impostor because it's like what we care about here is not just that it's going to get answers right, like get good test scores, like seem to be plausibly human when you talk about things that are in distribution. We want it to be able to go outside, like to do things that are creative, to be able to continue to learn, like to do to to get to the next level, including learn on its own and get to the next level. I mean, these are like the next frontiers for the field. If it's an impostor underneath the hood, then these kinds of things are are are going to hit a wall or become insanely expensive. It could be that you can always push through that wall, but the expenses just go up and up like crazy. Exponential, worse, I don't know what it means, but it could be something terrible. We might already be seeing that. Like, you know, the amount of money that we're spending here raises questions. Like, is it necessary? Like, does it have to cost this much in energy and in money? And so being an impostor is not a good thing even though the results look good. I mean, that's what it means to be an impostor. And we would like to know if that's the case. Well, we don't have direct hard evidence yet. What we have is some evidence from these pick breeder images that suggest there's a concern here, but we don't know how bad it is. And so it's just it's something that's, you know, good behooves the field to actually go off and figure this out. Speaking on the evidence, part of it, this goes into MechInterp research. People have been trying to dissect, like, how these models think and operate, for the longest time now. Right? And you it's, like, notoriously hard because they look into these, like, models and you have, like, stuff like polysemanticity where it's, like, 1 neuron is, like, activating via, like, adding circuits and also, like, black holes or something random. And it's, like, it's, like, it's so that's why it's so notorious. I think this is 1 of the fractured and tangled representations are 1 of the reasons why Beck and Terp is so hard is because how are you supposed to even interpret, something, that messy. Right? It's like it's just a hard problem. And this might I mean, I guess it begs the question, is neuroscience isn't easy either. Right? If you just try to look inside a human's brain, it's not like it's like smooth sailing when you look at a neuro human's brain either. But I guess, the question is, is there, like, a better organization for a neural network, which is also interpretable? Maybe it's not. I think we say in the paper, it's like a aspirational ideal. Maybe it's not possible, but it's a good, like, direction to go towards. Right? At least more going more unified and factored.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.