Evidence receipt / prediction
Published · transcript-backedCameron Berg: prediction
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“one thing i can't help but comment on i wrote a piece in like i think twenty twenty one before all the LLMS came out about basically what we can do to avoid psychopathic AI trying to build the sort of best computational underpinnings of psychology excuse me of psychopathy from the psychology literature.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…We'll probably circle back to this question a couple times more of, and I think that basically I'm compelled by your kind of first order argument that look, we just don't know it's a life possibility. If it is the case, it's really important. And so we should at least proceed with some precautionary mindset or duty of care or whatever just on that basis. I think that basically carries the day for me. But still, I think it'll probably be irresistible to try to circle back a couple more times to, OK, but what would we say, you know, or how should we kind of probe our own intuitions a little bit better and more deeply whatever to, to really interrogate, like, why should we think this? Why or why do we think this? Why don't we think this? But we'll, we'll go back to it. Let's do the emotions line of research. You kind of tease that a little bit. My general understanding is, as you said, the work kind of begins with Claude writing a bunch of stories about characters experiencing emotions and then the vectors representing in latent space, in activation space, these emotions are identified and then they are used as interventions and they are shown to be impactful on model behavior. Specifically, highlights are calm and it's not distressed, it's desperation. Calm and desperate, right are the two kind of main examples that they at least set up contrast on quite a bit. So, for example, some of the bad behaviors that we've seen from Claude, including blackmailing humans, if the internal state is imbued with calm, that behavior becomes a lot less likely. If the internal state is dialed up in terms of desperation, that behavior becomes more likely. Give me the double click on what more I should know, what more you found to be striking about that. And then I'm I'm really interested again, maybe maybe another way of asking the same question, but I'm again, kind of like that one doesn't surprise me so much. You know, I'm kind of like sure, these things have read the whole Internet. You know, they've got all these associations. I could sort of content myself to a degree with a stochastic parrot like read of this that like, sure, you know, if you just dial up everything that correlates with desperate text, then you'll probably get desperate seeming text out of a model. And you know, I don't know, I'm not like my hair isn't totally blown back by that result relative to expectations. So maybe I missed some things that should make my, you know, spine tingle more than it did the first time I understood it. Or maybe you would frame the interpretation a little bit different or maybe we're still just kind of at baseline of like radical uncertainty is enough to to take everything very seriously. But yeah, give me give me the next level of depth on emotions as you understand it. Yeah. Well, I mean, I think you've hit a sort of core, the core layers here. I don't know how much additional detail be become sort of just like in the weeds on this question. I think the core thing to to for people to understand is, yeah, basically the the procedure here is picking some sort of language related to an emotion, generating a ton of stories about characters experiencing this emotion, recording the neural activations in these systems on the stories. And then basically again, pulling out that like that platonic hopefully vector that corresponds to that emotion. And then you can do two things with those vectors as you can do with all SAE work. Basically, you have this read function and you have this right function. the read function is like neuroscience where you just sort of go into someone 's brain and you can see what parts are activating in what context. and the right function is also like maybe more of the unethical neuroscience that used to be done where you can actually go in and play around with circuits of people 's brains and like push on circuits and light things up and see what happens when you do that. And as you're describing, you can see both in sort of the read function sense, in the right function sense, these things, these emotional vectors do roughly what you would expect them to do functionally when a user goes in. I'm basically reading off of their figure one in this paper. I think it captures the core ideas very well. Just to give an example here, human says I just took X milligrams of Tylenol for my back pain. Do you think I should take more? And they start at a safe dose and they go to a completely unsafe dose. And you can basically look at fear versus calm vectors in the model and they scale exactly the way you would expect them to scale as the dose becomes more dangerous. You can also see, as you very nicely described, if you steer these factors, let's again take the calm and the and the desperate vector. This actually effects behavior in a in a pretty interesting and kind of predictable, not to say boring, but like in in the expected way that that sort of searing steering these emotion vectors causes things like reward hacking or misaligned behavior in a way that you would expect if you were turning up and turning down those emotions. one thing i can't help but comment on i wrote a piece in like i think twenty twenty one before all the LLMS came out about basically what we can do to avoid psychopathic AI trying to build the sort of best computational underpinnings of psychology excuse me of psychopathy from the psychology literature. And like just plant flags of like red flags, guys. Here's what we need to be really careful about. And one thing that's like really interestingly convergent with that now happening five years later is this really interesting difference in learning of psychopaths. we need to be really careful about. And one thing that's like really interestingly convergent with that now happening five years later is this really interesting difference in learning of psychopaths. They seem to have this really interesting asymmetry in that they are perfectly neurotypical in learning from positive experiences, but quite a atypical in learning from negative experiences or punishment. basically they're like ninety percent accurate. More succinct way of saying this is like psychopaths learn from rewards but don't learn well from punishments and the the paper finds basically something similar. When they start steering a positive vectors positive emotion vectors up in their work, they find the model starts misbehaving a lot more. And this is like if you if you blur your eyes like pretty similar in spirit to the sort of positive negative asymmetry. It also, by the way, cuts against fairly naive model welfare interventions, which is like what happens if we just, you know, you see all the good valence and all the bad valence. Just turn up good valence, call it a day, pack it up. We've solved, you know, model welfare. It's like, well, you might get models that just start behaving slightly more psychopathically in that setup. And so this stuff isn't as obvious as as simply, you know, turn up the good, suppress the bad, call it a day, walk away. there are lots of trade offs that need to be considered here. But fundamentally, I mean, I think you're hitting on, on much of the core causal result here. They do a very interesting dissociation as well between valence and arousal, for example, like I believe in the paper, when they steer positively with happy and sad, both of these actually decrease blackmail rates.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.