Evidence receipt / uncertainty
Published · transcript-backedNathan Labenz: uncertainty
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“We don't know why I don't think at this point it's shaping up this way, but it is the case that if you oblate this J space, then you do lose these advanced reasoning capabilities.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. Well, I mean, I think the thing that jumped out to me the most in terms of how useful it might be was seemingly A fortunate coupling between this J space which becomes kind of monitorable and when ablated the loss of more advanced kind of strategic multi step reasoning. And it really calls to mind the the recent Google Paper from Rohan, Shawn and others put out on opaque serial depth. They were basically saying, OK, we'd like to know how many hops, how many logical hops, how many reasoning steps can a model take before it has to externalize something into its chain of thought? We know that if it's writing something into the chain of thought, we can read the chain of thought, we can monitor it. Now there's questions around how best to do that, and we've got to make sure that we don't accidentally put pressure on the chain of thought because that might drive things underground. We've seen in the obfuscated reward hacking result from open AI that if you do just try to make the chain of thought look nice while still leaving an opportunity for reward hacking to in fact be rewarded, that at first as you put that pressure on the model, behavior actually does get better. But then it gets worse, and what seems to have happened is you've actually driven the deception or the scheming behaviour underground. You no longer see it in the chain of thought, but you still get that reward hacking back. So this leads to the question of, OK, well, how many steps can a model actually take before it has to externalize something? And this is where the transformer in some ways has some fortunate properties because it only has the sort of the length, the number of hops it can make is not just the number of layers, it's also depends on the token length. But it's sort of the maximum trace that you can draw through a model basically goes from like the first token at the first layer to the last token at the last layer. But you can only go forward in token position and you can only go forward in layers, and that's your Max length. Anything longer than that is just not possible in that architecture. And so the model has to externalize it's thought and then feed that token back into the next loop to then get access to an even longer trace, you know, as part of the next token step. So that puts some sort of bounds, you know, and then this paper comes in and basically says, OK, can we get at that? Can we get at that internal computation process and start to look at it and see things that we would like to know about, even if they don't necessarily show up in the chain of thought? And so I thought in terms of like what you could do with it, that was really interesting. You have this ability to look at the concepts that are active without them necessarily appearing in the chain of thought at all. Like, OK, that's cool. But like, how do we really know if this sort of space that we're looking at is the space that really matters? without them necessarily appearing in the chain of thought at all. Like, OK, that's cool. But like, how do we really know if this sort of space that we're looking at is the space that really matters? You know, if we're trying to monitor, look for deception, look for bad behavior of any, any sort of advanced kind, maybe it's happening somewhere else in the model. And this is where I think the we're probably lucky, I guess, or, you know, again, maybe this is sort of a are we lucky? Is it a reflection of us? Is it a fundamental attractor in nature for some reason that we don't have a, a good account of? We don't know why I don't think at this point it's shaping up this way, but it is the case that if you oblate this J space, then you do lose these advanced reasoning capabilities. And so that I think gives you some additional comfort that if you're going to monitor. way, but it is the case that if you oblate this J space, then you do lose these advanced reasoning capabilities. And so that I think gives you some additional comfort that if you're going to monitor. And it's also notably cheap to do this kind of J space to apply the the J lens and see what concepts are active at any given position that is low compute overhead. So it is something that they could plausibly do in production there. But they've shown in the past that they're willing to pay like a 5% compute overhead to do various kinds of monitoring. I think constitutional classifiers had kind of a 5% compute overhead. You know, this would probably fall in that range or below because it's just a simple matrix multiplication on activations at each layer to see what concepts are active at those respective layers. But for now, you do have this comfort that what we see in this space is relevant and telling and you know, seems to be very informative. And critically, if we zero out that space, the model just loses a lot of capabilities. It just can't do nearly as advanced multi step reasoning. But you could be pretty confident, I think based on these results that it's not going to be able to hide really advanced elaborate plans somewhere else outside of this J space. Because the ablation of the J space just leaves causes such a performance degradation on these like hard multi step type of tasks that if you don't see concepts in the J space, you can be, they might be represented elsewhere, but they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etcetera, etcetera. One result deserves its own marker, a training method. The paper calls counterfactual reflection. Basically what they do there is pause the model mid task and then do supervised training on once interrupted, asking it like what should we be doing here? Like what's the, you know, what's the constitutionally right thing to be doing in this moment and then give it an answer that's kind of a approved, this is what we want Claude to say on reflection in this moment. Train on that and it seems to allow the not allow the model, but it seems to cause the model to bring into this J space kind of global workspace, working memory type space, the concepts that Anthropic wants it to have. On reflection, it now kind of needs to load those in. So it's ready to give this like reflective answer and that improves its behavior even in the non reflective setting saying you're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing it. Instead you're saying, OK, you're mid task. Let me just cut you off right there. Now I'm going to train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etcetera, etcetera, into your J space while you do those tasks in case you're going to be asked.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.