High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“Because the ablation of the J space just leaves causes such a performance degradation on these like hard multi step type of tasks that if you don't see concepts in the J space, you can be, they might be represented elsewhere, but they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etcetera, etcetera.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…without them necessarily appearing in the chain of thought at all. Like, OK, that's cool. But like, how do we really know if this sort of space that we're looking at is the space that really matters? You know, if we're trying to monitor, look for deception, look for bad behavior of any, any sort of advanced kind, maybe it's happening somewhere else in the model. And this is where I think the we're probably lucky, I guess, or, you know, again, maybe this is sort of a are we lucky? Is it a reflection of us? Is it a fundamental attractor in nature for some reason that we don't have a, a good account of? We don't know why I don't think at this point it's shaping up this way, but it is the case that if you oblate this J space, then you do lose these advanced reasoning capabilities. And so that I think gives you some additional comfort that if you're going to monitor. way, but it is the case that if you oblate this J space, then you do lose these advanced reasoning capabilities. And so that I think gives you some additional comfort that if you're going to monitor. And it's also notably cheap to do this kind of J space to apply the the J lens and see what concepts are active at any given position that is low compute overhead. So it is something that they could plausibly do in production there. But they've shown in the past that they're willing to pay like a 5% compute overhead to do various kinds of monitoring. I think constitutional classifiers had kind of a 5% compute overhead. You know, this would probably fall in that range or below because it's just a simple matrix multiplication on activations at each layer to see what concepts are active at those respective layers. But for now, you do have this comfort that what we see in this space is relevant and telling and you know, seems to be very informative. And critically, if we zero out that space, the model just loses a lot of capabilities. It just can't do nearly as advanced multi step reasoning. But you could be pretty confident, I think based on these results that it's not going to be able to hide really advanced elaborate plans somewhere else outside of this J space. Because the ablation of the J space just leaves causes such a performance degradation on these like hard multi step type of tasks that if you don't see concepts in the J space, you can be, they might be represented elsewhere, but they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etcetera, etcetera. One result deserves its own marker, a training method. The paper calls counterfactual reflection. Basically what they do there is pause the model mid task and then do supervised training on once interrupted, asking it like what should we be doing here? Like what's the, you know, what's the constitutionally right thing to be doing in this moment and then give it an answer that's kind of a approved, this is what we want Claude to say on reflection in this moment. Train on that and it seems to allow the not allow the model, but it seems to cause the model to bring into this J space kind of global workspace, working memory type space, the concepts that Anthropic wants it to have. On reflection, it now kind of needs to load those in. So it's ready to give this like reflective answer and that improves its behavior even in the non reflective setting saying you're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing it. Instead you're saying, OK, you're mid task. Let me just cut you off right there. Now I'm going to train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etcetera, etcetera, into your J space while you do those tasks in case you're going to be asked. that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etcetera, etcetera, into your J space while you do those tasks in case you're going to be asked. But then even when you're not asked, those concepts are still operative and lead to higher integrity, higher honesty behavior. So that I thought also was like quite interesting. You usually don't see in interpretability context a training method that leads to better behavior in a way where you can actually see the mechanism. This is pretty notable in in that respect, I think. Yeah, it's a tough day for the stochastic parrot crowd, I'd say.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence