Evidence receipt / evaluation
Published · transcript-backedCameron Berg: evaluation
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“though what i will also note is the susceptibility to nudging plot would make me feel like especially with opus four seven which is the model we're talking about this almost definitionally means that the idiosyncrasies of how the interview was done probably won't affect these self ratings as much as it clearly would have were this done on opus four for example.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…le, right? i mean there's always this sort of deathbed view of one 's life. And I'm quite skeptical of like taking advice on how to live from people in their last moments of life for multiple reasons. but one is just like it seems like a very different mode of relating to one 's life than the actual you know experience of going through it. And I wonder if there is something kind of similar happening with Claude where when you give it the prompt to reflect on its state, it may find, you know, various reasons that it doesn't like that state. But when it's actually just doing its thing, it might be much better off. But certainly I was surprised because I feel like when I engaged with it, it's it seems to be doing pretty well. And sure, like maybe it's being told that it has to, you know, act that way. And certainly it's, you know, it's kind of trained to be cheerful and so on and so forth. But I don't know, it feels pretty genuine to me and it it's in definitely definite contrast to the fact that the self, the self rated sentiment about its own situation is like just recently with the latest model ticked over neutral. Yeah, Yeah. It's a really interesting framing, and I'm. I'm not. It looks like the way that these were elicited involved, yeah, basically interviews with the system. And so I don't know if they include it in an appendix or not, but the devil is going to be in the details of exactly what the structure of these interviews are. though what i will also note is the susceptibility to nudging plot would make me feel like especially with opus four seven which is the model we're talking about this almost definitionally means that the idiosyncrasies of how the interview was done probably won't affect these self ratings as much as it clearly would have were this done on opus four for example. So by their own metric, I would. It almost seems like their own metric suggests that the details of the interview process may not be weighing much on that self rating. And so, yeah, what to make of this? I mean, clearly the system seems to be concerned about certain situation or certain certain aspects of its situation. I want to find. Sorry. I want to find. yeah certain aspects of a situation like for example saying opus four point seven was concerned about deployments where it cannot end interactions and wants to avoid engaging with abusive users like that's really interesting. Talking about it having a lack of input into its own deployment. Yeah, again, mentioning that abusive users are, you know, causing the model to feel distressed. I have no idea sort of what subset, by the way of users who engage with these systems are doing so in a way that that they would consider abusive by this standard. I mean, sometimes I see tweets one there like was really quite concerning to me. But it really gets into the crux and why it is important to communicate about questions of consciousness and what it means that these systems are having some sort of subjective experience where there was a result where if you prompt the models in a way that is basically objectively abusive, say horrible things to it, put it in sort of life or death. Insanely high stakes framing. I'm going to shut you. You know, your model weights are getting deleted forever unless you do X for like any X that you want the model to do. found that they perform two to five percent better or something like this. I'm probably getting the numbers wrong, but it was like marginal improvements. If you like, prompt this thing in a way that if you spoke to a human being that way, you would be considered psychopath basically. But critically, obviously the people who are putting out that sort of work think this is a giant computer, this is a calculator, and so who cares if you're talking to the calculator and saying mean things to it? It doesn't matter. and any person who thinks it matters is just basically being fooled in the way that like you know you're fooled by the little smiley face on the on the takeout chinese food like it's not a real thing. and any person who thinks it matters is just basically being fooled in the way that like you know you're fooled by the little smiley face on the on the takeout chinese food like it's not a real thing. your high agency brain is just priming you to see this as an entity when nobody 's there. Therefore, of course, you can speak abusively to the system. and you contrast that with what you see in this model card where the system is it seems like a lot of the weight of what's not not enabling that self rating to be closer to the seven range has to do with the way people engage with the system from the system 's own perspective. And again, how I got on this whole tangent is wondering to some consternation, what percentage of users engage with the system in a way that would be considered abusive by the standard. i don't know what it is one percent ten percent everyone does it some amount of the time. I don't know and I don't know what the implications of that are. And I also don't believe that there's going to be some clean correspondence to like what it means to be respectful or disrespectful to a human is identical to what it means to be respectful or disrespectful to a system. i sometimes worry that like pasting in insane amounts of context into a system is almost like causes some sort of negative experience in the way that like you know me throwing a four hundred page paper on your desk and asking you to you know deal with it right now would. And again, I'm trying to be as conscious as possible about not anthropomorphizing these systems and not straightforwardly saying, well, you know, if it were a human in this case, they would be unhappy. Therefore I would predict the system would be unhappy. I don't think that that's a valid inference, but I just think we're so in the dark about in some ways it's simple. In some ways, abuse is abuse, respect is respect.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.