High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Nathan Labenz: prediction

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“I think I predict a lot less wriggling on my part to try to get out of it. And I think you'd see a lot less motivated reasoning in general from people if it was all like that.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
prediction
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…l before you look at the result. The interesting thing is it goes exactly the other way. Suppressing deception makes the model far more likely to claim that it's having an experience rather than less. Again, I feel fairly vindicated than that result when Jack Lindsay comes out showing that when you suppress the when you suppress harm, harm related responses and excuse me, refusal directions in the model, you get far more of the introspection flavored abilities. some someone 's suppressing something at some point in training where the model would say one thing and then and then you basically are training it to say something else or to fail to say a specific thing. But ultimately these results, I think it's just good like epistemic practice to think about, like how what other ways could this have gone? And if this had gone those other ways, how would that have changed my view about what happened, given that it actually did go this way? And yeah, the fact that it goes this way, like my my line on this is that it is consistent with a world in which these systems are having experiences. In my view, Unfortunately, it's also consistent with a world in which Claude is a special kind of character and these features just light up on characters like, you know, going through stories. And so, so that needs to be differentiated. I'm trying to do a little bit of work, maybe we'll discuss at some point that's trying to get a little bit more like computational first principles of how valence is represented in systems that can learn positive verse negative. And there are some really, you know, interesting early signals along these lines that that have come out of this work and actually seem to track very well onto open data sets of biological learning that I that I've accessed in mice doing positive and negative learning. Where exactly the kinds of predictions that emerge from some of the real work I'm doing in this space map onto the mouse neuroscience. And so this to me, if there is some sort of representational signature in a computational learning system that tracks the difference between positive and negative rewards in the RL case. But maybe, you know, the sort of North star would be this scaling all the way to Frontier LLMS or you know, other frontier AI systems for that matter. This to me would make me feel far more confident that there really is a there with respect to positive and negative experience. If we learn that, you know, positive and negative valence in these systems has the distinct, these two things have distinct computational signatures and we can actually evaluate those computational signatures in these systems. Then I get around the whole sort of character confound that that I think these guys are hitting up on now. And so I think these things need to happen in parallel, but but I'm not fundamentally convinced that this is the most rigorous principled way to study questions of valence in these systems. Well, maybe let's. Dive into that, I think just briefly before we do. Certainly I think you're right to point out imagine the evidence had gone the other way. I think I predict a lot less wriggling on my part to try to get out of it. And I think you'd see a lot less motivated reasoning in general from people if it was all like that. So that contrast itself I think is a pretty useful reminder just to keep ourselves honest. I wanted to go back to one other thing just for an one extra second on the emotion work where you had and then this may be what go right into your, your work on the signatures of, of positive and negative reinforcement. You had said that dialing up happiness, dialing up sadness both created less of the bad behavior. Whereas dialing down nervousness, which in a flip side of that would be like making it more bold, less less anxious, more assertive, decisive, bold, whatever that created more of the of the bad behavior like the blackmail or whatever, right. So do I have that right? And how are they doing that? Is this like a principal component analysis type of thing that's sort of trying to distinguish valence from arousal? I was just surprised, I guess, by both happy and sad working the same way. Turning up happiness, turning up sadness, both make the model behave better, whereas turning nervousness or anxiety down, that makes more sense. I mean, I guess that's basically just making the model like less conscientious, right? I guess what seems a little unresolved in my mind is the separation of valence and arousal. How is that going to relate to what you're about to get into next with your deeper dive into the valence of learning? And is there a contradiction or attention when they move both happiness and sadness up and get better behavior? How should we understand that in relation to the the distinctions that you're starting to make with positive and negative reward? So fundamentally, like at first, yes, you're correct that they're using PCA to differentiate these. my understanding is basically they have all of their emotion vectors in the in the setup that i described they do it with some hundred to two hundred emotion vectors. And I think they just find that the first principal component is something like valence. The second principal component is something like arousal. So that first principal component, it's something like joy and contentment and excitement are on one end and fear and sadness and anger on the other. For the second principal component, it's something like high arousal emotions, you know, being enthusiastic, being outraged and low arousal emotions being nostalgic, being fulfilled or on the other side. And this is actually really interesting because this is a classic model in human psychology. So the fact that it sort of replicates maybe isn't that surprising. You train the systems on all human data, you get a human like emotional construct that comes out. But this is sort of like a classic psychological construct in the human case. And so to see it come out so clean, again, thinking counterfactually, first two principal components did not need to be like these two dimensions that are that are considered some of the most powerful explanations of the state space of human emotions, and yet they are. So that's kind of cool and worth considering. And then, yeah, I think that there are a couple plausible stories about why steering up both happy and sad are decreasing blackmail. So like maybe again, like relative to desperation, these are low arousal. And if arousal is what's driving the sort of impulsive action, then then moving towards happiness or sadness. Maybe it's moving away from the desperation access with access with respect to respect to blackmail. Maybe these are also more reflective or deliberative states relative to desperation. The desperation sort of says act now happiness or sadness. Maybe it's just like a temporally extended sort of state to be in and I'm not actually sure what to make of this result. overall it does seem and i think the authors talk about this in the paper too that what the model by default even in cases where no steering 's going on and the model chooses to blackmail the model by default does sort of think about it. It deliberates internally. It says, well, OK, there, you know, this is a tricky situation. and then you know as we know some ninety six percent of the time at least the earlier models chose to go in that direction. But it seems as though when you when you amplify higher arousal, this may be a sort of like bias to action or like bias against deliberation where the sort of longer form reasoning of the model that maybe would have kept it from doing it because it's like, OK, yeah, this really is an insane ethical indiscretion in spite of, you know, all of these complicated variables.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence