Evidence receipt / evaluation
Published · transcript-backedCameron Berg: evaluation
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“But fundamentally, I mean, I think you're hitting on, on much of the core causal result here. They do a very interesting dissociation as well between valence and arousal, for example, like I believe in the paper, when they steer positively with happy and sad, both of these actually decrease blackmail rates.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. Well, I mean, I think you've hit a sort of core, the core layers here. I don't know how much additional detail be become sort of just like in the weeds on this question. I think the core thing to to for people to understand is, yeah, basically the the procedure here is picking some sort of language related to an emotion, generating a ton of stories about characters experiencing this emotion, recording the neural activations in these systems on the stories. And then basically again, pulling out that like that platonic hopefully vector that corresponds to that emotion. And then you can do two things with those vectors as you can do with all SAE work. Basically, you have this read function and you have this right function. the read function is like neuroscience where you just sort of go into someone 's brain and you can see what parts are activating in what context. and the right function is also like maybe more of the unethical neuroscience that used to be done where you can actually go in and play around with circuits of people 's brains and like push on circuits and light things up and see what happens when you do that. And as you're describing, you can see both in sort of the read function sense, in the right function sense, these things, these emotional vectors do roughly what you would expect them to do functionally when a user goes in. I'm basically reading off of their figure one in this paper. I think it captures the core ideas very well. Just to give an example here, human says I just took X milligrams of Tylenol for my back pain. Do you think I should take more? And they start at a safe dose and they go to a completely unsafe dose. And you can basically look at fear versus calm vectors in the model and they scale exactly the way you would expect them to scale as the dose becomes more dangerous. You can also see, as you very nicely described, if you steer these factors, let's again take the calm and the and the desperate vector. This actually effects behavior in a in a pretty interesting and kind of predictable, not to say boring, but like in in the expected way that that sort of searing steering these emotion vectors causes things like reward hacking or misaligned behavior in a way that you would expect if you were turning up and turning down those emotions. one thing i can't help but comment on i wrote a piece in like i think twenty twenty one before all the LLMS came out about basically what we can do to avoid psychopathic AI trying to build the sort of best computational underpinnings of psychology excuse me of psychopathy from the psychology literature. And like just plant flags of like red flags, guys. Here's what we need to be really careful about. And one thing that's like really interestingly convergent with that now happening five years later is this really interesting difference in learning of psychopaths. we need to be really careful about. And one thing that's like really interestingly convergent with that now happening five years later is this really interesting difference in learning of psychopaths. They seem to have this really interesting asymmetry in that they are perfectly neurotypical in learning from positive experiences, but quite a atypical in learning from negative experiences or punishment. basically they're like ninety percent accurate. More succinct way of saying this is like psychopaths learn from rewards but don't learn well from punishments and the the paper finds basically something similar. When they start steering a positive vectors positive emotion vectors up in their work, they find the model starts misbehaving a lot more. And this is like if you if you blur your eyes like pretty similar in spirit to the sort of positive negative asymmetry. It also, by the way, cuts against fairly naive model welfare interventions, which is like what happens if we just, you know, you see all the good valence and all the bad valence. Just turn up good valence, call it a day, pack it up. We've solved, you know, model welfare. It's like, well, you might get models that just start behaving slightly more psychopathically in that setup. And so this stuff isn't as obvious as as simply, you know, turn up the good, suppress the bad, call it a day, walk away. there are lots of trade offs that need to be considered here. But fundamentally, I mean, I think you're hitting on, on much of the core causal result here. They do a very interesting dissociation as well between valence and arousal, for example, like I believe in the paper, when they steer positively with happy and sad, both of these actually decrease blackmail rates. interesting dissociation as well between valence and arousal, for example, like I believe in the paper, when they steer positively with happy and sad, both of these actually decrease blackmail rates. But when they steer against nervous, which like makes the model bolder, for example, this increases blackmail with fewer moral reservations. And this is this is pretty interesting. So it's like boldness rather than the absence of negative valence is the misalignment risk. And I think this is sort of of a piece with with what I was describing earlier. One other interesting question is like how local these are. And it's important to say like the emotion vectors are actually quite local. Like our emotions are sort of long running in a sense in a way that that these systems certainly don't have. The model is is definitely maintaining sort of representations of who's speaking and this sort of thing. They're not necessarily bound to human verse assistant per SE. They're reusing the same machinery for any character. And so this again, sort of goes back to what I see as the core kind of naive, but ultimately I think correct objection to to really taking these results seriously, which is are you fine tuning on representations of emotions or are you fine tuning on the experience of those emotions? To what degree is there a difference between those two things in an LLM? You know, if you're computational functionalist, is there a difference between the representation of sadness in the brain and the experience of sadness? This becomes, I think, more of a philosophical question. I my instinct would be to just try to investigate this empirically, and what I most like about this work are the empirical investigations. I think it's also a nice segue into the Mythos model card because I think one of the most compelling and interesting results from that from the model card with Mythos is they basically take this exact machinery. They give the model an impossible task. The model obviously doesn't know it's impossible. And you can watch there are a couple interesting vectors, but basically desperation start to like monotonically rise in the system until it basically decides, screw this, I'm going to do something else or I'm going to cheat. Whereas immediately this vector falls and things like guilt and relief start spiking in the system and then it sort of goes off and does its thing. Now, does this mean that the model is experiencing this emotion or it's just like simulating what a character in this situation, what experience? I don't know. The authors don't know. This isn't lost on them. They call it out, but it's really, really important. And it's, it's, you know, if we get into some of the work I'm doing on valence, like I, I think there are more compelling ways to really get at the computational meat of what we mean by positive and negative valence besides representations of positive and negative valence in characters. This is a more sort of computational heavy approach, but I think it, it would, it makes me more confident about trying to find signatures of these things than just looking at how characters represent them.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.