High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Nathan Labenz: uncertainty

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“Is it convergent evolution or is it a reflection of, of how we think somehow encoded in the data that it's then reverse engineering our structure from the shadow of that structure as it's encoded in text? It's a, it's a very interesting question and I don't know that the paper really has anything to say about that yet, but there's certainly going to be a lot of future work, I think, downstream of this one.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
uncertainty
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Before the big picture, what this paper did to a prior of mine, starting with their announcement video. The video is beautiful. I, I enjoyed the video quite a bit. You know, it set off my alarm bells a little bit in terms of how much they're really embracing anthropomorphizing the models at this point. I used to say beware overly anthropomorphizing. You know, remember that these are these things are so alien and we shouldn't assume that the way that we work is the way that they work. And I have to say that has come due for some significant revision. People that have embraced anthropomorphizing, I think have got quite a lot of mileage out of it. I do something, it's obviously something to be really careful about and as much depth and detail as there is in this research, it's like easy to, you know, maybe get carried away with it and forget, you know that there are a lot of caveats as well and there's a lot of things where like it doesn't always work. This was kind of my big concern with the tracing large language model thoughts paper. There's just like a lot of residuals and there's a lot of error correction terms along the way that they use to make that thing work. And when you have the kind of zoomed out trace view and you're like, oh, OK, so this is how it works. Like this gets loaded in and these two features interact and they, you know, kick out this third feature. And that's how we get our answer. It's easy to forget just how much kind of fuzziness and not full story there was along the way toward that stylized account. So I think it's going to be very important for everybody from the researchers that entropic to the public to kind of hold two thoughts in mind at the same time. All the caveats trying to hold those in mind, it is still a big update toward anthropomorphizing being a valid and in many cases productive approach for thinking about language models. I would not have expected the cognitive machinery of a large language model to look so similar structurally to the human version, as best we understand it, it seems to. And I wouldn't have expected that theories of human cognition would motivate so many good experiments on LLMSI just would have expected the show Goth to be far more alien and to have mechanisms, you know, far more different than our own. And it leaves me now wondering to what degree is this a sort of natural result of physics? You know, is there some sort of, you know, when you're trying to do cognition under budgetary constraints? Are these mechanisms just the natural mechanisms that emerge? Or is this in some way a reflection of us in the, in the data, you know, is in other words, if you were somehow to train an AI without basing it on so much human data, would we see similar structures emerge? Or would it would we go back to a more, you know, alien hypothesis where they're just totally, totally different? And, you know, there's little in the way of analogical structures between the two processes. ld it would we go back to a more, you know, alien hypothesis where they're just totally, totally different? And, you know, there's little in the way of analogical structures between the two processes. And I, I don't have a great intuition for that at this point, but it definitely has me asking the question because they're, they are over and over again. It seems to be coming in that they're much more structurally similar to us that I would have guessed. And, and yeah, is that something that nature just finds as kind of a convergent solution? Is it convergent evolution or is it a reflection of, of how we think somehow encoded in the data that it's then reverse engineering our structure from the shadow of that structure as it's encoded in text? It's a, it's a very interesting question and I don't know that the paper really has anything to say about that yet, but there's certainly going to be a lot of future work, I think, downstream of this one. So I saw something from Daniel Cocatello, who was like everybody, you know, impressed, but also said, I think his tweet was just a few dozen more advances like this and we might actually be able to make the AIS really safe. Daniel Cocatello, who was like everybody, you know, impressed, but also said, I think his tweet was just a few dozen more advances like this and we might actually be able to make the AIS really safe. I am, I think, a little bit more optimistic than that. What is my track record as a super forecaster here or as a borderline super forecaster for my results on the original Tetlock thing years and years ago? I think my track record is that I probably tend to underestimate, you know, how many more breakthroughs will be needed, you know, for anything. So my analysis should be colored with that bias or or, you know, awareness of my possible weakness in that regard. But another way to stay, say, my intuition around how much room is there left to hide, is that I don't feel like we need dozens more insights of this scale to get to a point where we might actually be able to keep this thing on the rails. It's going to be tricky. I don't want to make it sound easier than it is, but I don't know that we need all of the AIS to be aligned right or that it's such a big problem if somebody out there somewhere does something problematic. Now, it could be very problematic if they create the thing that launches the, you know, the next pandemic and literally kills us all with an engineered pathogen or whatever. So there is some mechanism where that could go super, super bad. We're going to need to like harden the world, the pandemics, no doubt about that, for all sorts of reasons. And AI probably being the the biggest. But I also remember this thing that Zuckerberg said once that I thought was pretty compelling. He basically said, you know, we deal with scammers and spammers all the time. And the big advantage we have is we have all the compute, we have all the resources, you know, but at a systemic level, we're just way bigger, way better, way more sophisticated than them. And we can also have like really big institutional developers like Anthropic and like Google and, and Open AI that like, I think between those three companies, you know, they're going to have something like pushing half of global compute to, to work with. So if they do a really good job on this kind of stuff, if they, you know, a, a handful of these things feel like maybe they could be enough. Like maybe there maybe it cuts the light in space in enough different ways. Maybe we have enough lenses on it that we really can come to a pretty strong conclusion. Like, hey, there really isn't much space left in this thing to hide. And yeah, there might be some bad biases or bad attitudes or bad, whatever. But if we can, if we can put upper bounds on how schemy the model can really be, because we can look at it through this and you know, a handful of other different lenses and and have pretty reliable takes on that. I'm more than ever before, I feel like, and this has been growing. This is not the first positive update, but more than ever before, I'm like, you can maybe imagine a super intelligent quad that we could actually have enough insight into to be pretty confident that it's actually trying to do the right thing for us.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence