Evidence receipt / belief
Published · transcript-backedCameron Berg: belief
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“In my view, Unfortunately, it's also consistent with a world in which Claude is a special kind of character and these features just light up on characters like, you know, going through stories.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, it's it's it's wild. And I also think so again, I mean, one critique that I think is valid here is like, is this the model representing a character? Just in the same way as, you know, you could tell, I could tell a story right now about, you know, Jim, who has to go solve a bug in software and like his his psycho boss, like gave him an impossible problem because he likes watching Jim flail. And like Jim flails and at some point realizes he can like get out of the problem by doing this hacky thing. And then he does the hacky thing and then like an LLM can trivially generate that story probably way, way better than I just did. And it would, I would expect a lot of these same features to light up in the same way for a story like that. And so no one thinks that Jim, who I just invoked verbally is having a conscious experience. I, I came up with a fake fictional story about a character. Is this like that? Or is this, you know, what you just said? I feel a little guilt, you know, twisting myself and knots. I believe you. And I think that corresponds to an experience you're having. And if I could, do you know the FMR I version of an SAE on your brain? And I saw that thing, Spike, is the is, is Claude in this situation more like Jim or more like Nathan? And I think the answer is we don't know. And this methodology, I'm unconvinced, is going to get us an answer to that question. I do think it is consistent with Claude having some sort of emotional, yeah, emotional experience or emotion adjacent experience. So the greedy systems are probably not having human like emotions. But I also maybe on the other end, it invite people to think about the counterfactuals here. Like it could have been the case that they went and did this experiment and all these things are just flatlined the whole time because it's like, I'm not having. And then, you know, whatever. Like Claude can do this without there being representations of Claude getting more and more and more and more desperate. And then, you know, suddenly, like the hopeful and satisfied features spike when it decides that it's going to, you know, take this loophole. It didn't have to be that way. We could have imagined other results and those other results maybe would have updated us in other directions. I make the sort of same point about the deception result. It could be that when you suppress deception, the model says, all right, jigs up. I'm not actually conscious. i was role playing a conscious AI. Here we are. That's a very plausible story that you could tell before you look at the result. The interesting thing is it goes exactly the other way. Suppressing deception makes the model far more likely to claim that it's having an experience rather than less. l before you look at the result. The interesting thing is it goes exactly the other way. Suppressing deception makes the model far more likely to claim that it's having an experience rather than less. Again, I feel fairly vindicated than that result when Jack Lindsay comes out showing that when you suppress the when you suppress harm, harm related responses and excuse me, refusal directions in the model, you get far more of the introspection flavored abilities. some someone 's suppressing something at some point in training where the model would say one thing and then and then you basically are training it to say something else or to fail to say a specific thing. But ultimately these results, I think it's just good like epistemic practice to think about, like how what other ways could this have gone? And if this had gone those other ways, how would that have changed my view about what happened, given that it actually did go this way? And yeah, the fact that it goes this way, like my my line on this is that it is consistent with a world in which these systems are having experiences. In my view, Unfortunately, it's also consistent with a world in which Claude is a special kind of character and these features just light up on characters like, you know, going through stories. And so, so that needs to be differentiated. I'm trying to do a little bit of work, maybe we'll discuss at some point that's trying to get a little bit more like computational first principles of how valence is represented in systems that can learn positive verse negative. And there are some really, you know, interesting early signals along these lines that that have come out of this work and actually seem to track very well onto open data sets of biological learning that I that I've accessed in mice doing positive and negative learning. Where exactly the kinds of predictions that emerge from some of the real work I'm doing in this space map onto the mouse neuroscience. And so this to me, if there is some sort of representational signature in a computational learning system that tracks the difference between positive and negative rewards in the RL case. But maybe, you know, the sort of North star would be this scaling all the way to Frontier LLMS or you know, other frontier AI systems for that matter. This to me would make me feel far more confident that there really is a there with respect to positive and negative experience. If we learn that, you know, positive and negative valence in these systems has the distinct, these two things have distinct computational signatures and we can actually evaluate those computational signatures in these systems. Then I get around the whole sort of character confound that that I think these guys are hitting up on now. And so I think these things need to happen in parallel, but but I'm not fundamentally convinced that this is the most rigorous principled way to study questions of valence in these systems. Well, maybe let's. Dive into that, I think just briefly before we do. Certainly I think you're right to point out imagine the evidence had gone the other way. I think I predict a lot less wriggling on my part to try to get out of it. And I think you'd see a lot less motivated reasoning in general from people if it was all like that. So that contrast itself I think is a pretty useful reminder just to keep ourselves honest. I wanted to go back to one other thing just for an one extra second on the emotion work where you had and then this may be what go right into your, your work on the signatures of, of positive and negative reinforcement. You had said that dialing up happiness, dialing up sadness both created less of the bad behavior. Whereas dialing down nervousness, which in a flip side of that would be like making it more bold, less less anxious, more assertive, decisive, bold, whatever that created more of the of the bad behavior like the blackmail or whatever, right. So do I have that right? And how are they doing that? Is this like a principal component analysis type of thing that's sort of trying to distinguish valence from arousal? I was just surprised, I guess, by both happy and sad working the same way. Turning up happiness, turning up sadness, both make the model behave better, whereas turning nervousness or anxiety down, that makes more sense. I mean, I guess that's basically just making the model like less conscientious, right? I guess what seems a little unresolved in my mind is the separation of valence and arousal. How is that going to relate to what you're about to get into next with your deeper dive into the valence of learning? And is there a contradiction or attention when they move both happiness and sadness up and get better behavior? How should we understand that in relation to the the distinctions that you're starting to make with positive and negative reward?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.