Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“You usually don't see in interpretability context a training method that leads to better behavior in a way where you can actually see the mechanism. This is pretty notable in in that respect, I think.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…way, but it is the case that if you oblate this J space, then you do lose these advanced reasoning capabilities. And so that I think gives you some additional comfort that if you're going to monitor. And it's also notably cheap to do this kind of J space to apply the the J lens and see what concepts are active at any given position that is low compute overhead. So it is something that they could plausibly do in production there. But they've shown in the past that they're willing to pay like a 5% compute overhead to do various kinds of monitoring. I think constitutional classifiers had kind of a 5% compute overhead. You know, this would probably fall in that range or below because it's just a simple matrix multiplication on activations at each layer to see what concepts are active at those respective layers. But for now, you do have this comfort that what we see in this space is relevant and telling and you know, seems to be very informative. And critically, if we zero out that space, the model just loses a lot of capabilities. It just can't do nearly as advanced multi step reasoning. But you could be pretty confident, I think based on these results that it's not going to be able to hide really advanced elaborate plans somewhere else outside of this J space. Because the ablation of the J space just leaves causes such a performance degradation on these like hard multi step type of tasks that if you don't see concepts in the J space, you can be, they might be represented elsewhere, but they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etcetera, etcetera. One result deserves its own marker, a training method. The paper calls counterfactual reflection. Basically what they do there is pause the model mid task and then do supervised training on once interrupted, asking it like what should we be doing here? Like what's the, you know, what's the constitutionally right thing to be doing in this moment and then give it an answer that's kind of a approved, this is what we want Claude to say on reflection in this moment. Train on that and it seems to allow the not allow the model, but it seems to cause the model to bring into this J space kind of global workspace, working memory type space, the concepts that Anthropic wants it to have. On reflection, it now kind of needs to load those in. So it's ready to give this like reflective answer and that improves its behavior even in the non reflective setting saying you're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing it. Instead you're saying, OK, you're mid task. Let me just cut you off right there. Now I'm going to train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etcetera, etcetera, into your J space while you do those tasks in case you're going to be asked. that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etcetera, etcetera, into your J space while you do those tasks in case you're going to be asked. But then even when you're not asked, those concepts are still operative and lead to higher integrity, higher honesty behavior. So that I thought also was like quite interesting. You usually don't see in interpretability context a training method that leads to better behavior in a way where you can actually see the mechanism. This is pretty notable in in that respect, I think. Yeah, it's a tough day for the stochastic parrot crowd, I'd say. And because you can see this and you know, these models are billions of parameters at this point, and they did it also on 4.5 Sonnet. 4.5 Sonnet is actually a pretty recent model. It's not. It's just like seven months, 8 months since 4.5 Sonnet and Sonnet is a very capable model. So it is it is a fairly large model to apply this thing to. I'm not sure if you saw the Neil Nanda commentary, he they applied it on a QNQ N27B. So that that is also a pretty, pretty advanced advanced model.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.