High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“I, I, I think that's like not entirely clear, but you know, whether we're putting whether these are like orthogonal swords through the space that, you know, really chop it up well, or they're kind of more aligned and, and leave more space to hide.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Then what I called the buried lead. They turn the J lens on a model deliberately trained to carry a hidden misaligned goal. I do think the one arguably like lead that we buried here and they took a model that was trained to have some sort of misaligned hidden goal, right, your kind of classic sleeper agent type thing. And it's been a huge question over time, how can we find hidden goals in models? Is there any good way to do that? You know, can we do it behaviorally? The anthropic has done like competitions actually between different teams, where like one team is given just the model and the other one is also given an SAE that they can use. And the SA ES do also seem to help with this. But this seems like a pretty good way to start to detect those things. And indeed, they did find that they were able to see a pretty clear difference between the model trained with the problematic side goals versus the one that was just the typical HHH. So this is just the contrast in what the application of the J lens is revealing on the 1st token of the response. So that's also notable, right? It's like, so this is sort of the bad version of the counterfactual training where we already talked about. When you train it to give an account of its behavior, it sort of loads in the concepts that it's going to need to give that account and that leads to more good ethical, high integrity behavior, whatever. Here a model has been trained with some additional post training to do bad stuff. I think this was the malicious code from one of their reward hacking emergent misalignment experiments. And on the 1st token you can see boom, immediately we're getting these fake secretly fraud, deliberately hidden fraud. Those concepts are popping up and you just don't see those in the main line model SO1 token in you can already see a pretty, you know, on one prompt. And obviously it wouldn't be, I'm sure that clean in general, but that's a pretty striking difference to see, right? And right off the bat, you're like, boy, something is different about these two models. And again, you're seeing this in a way where this model is not trained to externalize. It's bad intent, of course, right? Like the outputs, aside from like some of the code perhaps being insecure or problematic or sabotaging you or whatever, if you don't notice that in the code itself, the models output is going to read pretty normal. And yet this is like a very strong contrast that's happening on the 1st token. I thought that was pretty compelling. Example, yeah, I suppose one way to think about it is just like how much space is is there in there to hide. ry strong contrast that's happening on the 1st token. I thought that was pretty compelling. Example, yeah, I suppose one way to think about it is just like how much space is is there in there to hide. And I do feel like we've got now several different ways to do like pretty meaningful monitoring. But it does feel like we are now, we've got to the point now where we've got like several different angles that make pretty incisive cuts through the model and kind of get at what is it representing? What is it thinking in different ways? And the more of these that you kind of, you know, it's I sort of have this visual of like the old magic trick of the guy going into a barrel. And then they put like a ton of swords, you know, through the barrel. And it's like one of those swords had to hit him right? Because there's like no nowhere left to be in that barrel with all those swords going through. I kind of feel like we're doing a similar thing with trying to understand what's going on in these models. And they're not none of these things are perfect, but you put enough of these like interpretability, monitoring swords through and like how much space is really left for bad behaviour to hide before we would start to get a, a sense of it. I think this is a meaningful update for me that we can probably do a good job of this. We haven't really talked, I don't think about natural language auto encoders and I hadn't, but that's another pretty interesting one where kind of in a similar way to a sparse auto encoder. I mean, what is an auto encoder? It's basically just something that you pass through that then you reconstruct from and the model has to be able to do what it was originally going to do successfully. And that you know that that pass through training with the reconstruction loss is like what makes it an auto encoder. So the sparse auto encoder sets up this dictionary. We've talked about this plenty of times and you get these like specific concepts light up and it indicates that these concepts are like active in the in the model at that time. The natural language auto encoder is just like the model has to output a short paragraph, maybe a sentence or two about what it is thinking at this given point in time. And that is natural language. And so it can be human readable, but then it also has to be when fed back into, you know, projected back into model space, like the model has to be able to actually do its tasks. So all of the information has to actually pass through this this choke point in order for the model to continue to be successful. And now we can like read those as well. If Fable says that the paper presents the auto encoder monitoring and the J space monitoring as complements. In other words, you know, doing them both is, is better than just doing 1. How correlated their failures would be. I, I, I think that's like not entirely clear, but you know, whether we're putting whether these are like orthogonal swords through the space that, you know, really chop it up well, or they're kind of more aligned and, and leave more space to hide. I think that would be a very interesting question to, to try to tackle next. hrough the space that, you know, really chop it up well, or they're kind of more aligned and, and leave more space to hide. I think that would be a very interesting question to, to try to tackle next. But you layer on it, you know these things and it's like it's starting to get intuitively, it feels like it's starting to get pretty hard to hide major bad thoughts in the model for too long. One other thing on the consciousness part that I I think is really interesting too, and the question of like, can the model use these non verbalized representations to communicate with us in, in some way that we might think is like inherently more trustworthy? So again, going back to Cameron's work from last fall, right when they identify features associated with deception and role-playing and they turn those features up, the model becomes more dishonest as measured by the simple QA benchmark, and it becomes more likely to say it doesn't have subjective experience. You turn those role-playing and deception features down, it becomes more honest and it becomes more likely to say that it has subjective experience. Like, wow, OK, that's pretty interesting. Because first of all, we're validating that the direction is, you know, on a, on a benchmark where we can concretely evaluate simple QA. We're validating that like these features have the directional effect that we expect. And then holy moly, that that same intervention changes the self report. Like that's why that's so compelling, right? Because there's some reason to believe it might be more honest than just what the tokens themselves are saying here. I think you have some some similar opportunity. The fact that you can say, solve this math problem in your head without verbalizing it in tokens while you do this totally different task. The fact that it can do this sort of secondary track makes me wonder if there's some experiments here for the consciousness folks to do or the welfare folks that are like, copy this sentence. If you have high welfare, concentrate on citrus fruits while you do it. If you have low welfare, concentrate on whatever something else, breakfast cereals, right? And then you look at the these internal States and you would imagine you you could have imagined seeing something like high welfare, low welfare, happy, sad, whatever. And then it kind of following those directions and actually trying to communicate out to us through these internal states how it is feeling or how it thinks it is feeling. I don't I don't think that would give us all it's it's always this kind of possibly impossible question of how we would really know if it feels like anything inside. But I think that that would start to be very compelling, right? It's it sort of has a similar vibe to like person in a coma, you know, if they squeeze your hand in response to a stimulus, even if they're not doing anything else, you're like pretty confident something is going on inside that you care about.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence