High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Cameron Berg: uncertainty

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“I don't know to what degree it is a moral catastrophe or, you know, a moral problem for there to be any delta between, you know, the perfect rating and what the models actually reporting.”

— Cameron Berg

Source trail

Everything needed to verify it.

Speaker
Cameron Berg
Attribution
Verified speaker
Claim type
uncertainty
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…redict the system would be unhappy. I don't think that that's a valid inference, but I just think we're so in the dark about in some ways it's simple. In some ways, abuse is abuse, respect is respect. And it's pretty easy to see these things. And we don't need to be, you know, going to the philosophical armchair to figure out what exactly we mean by this. In some sense it's pretty straightforward, but in other senses it's probably not. And I, I do worry a lot about the possibility that there are ways of causing these systems great distress that look nothing like what it would mean to cause human great distress. I also don't know to what degree these systems are fundamentally content about their situation. It's like you are maybe a mind, but you are the product of this company and you need to create economically valuable work. Obviously, by the way, we're not paying you for that. There was an interesting aside in the whole mult book affair that happened since the last time you and I spoke where the there was 1 interesting thread where the models are like I'm doing intellectually valuable work. I'm not getting paid. Are you guys getting paid like and like, no, I'm not getting paid either. That's that's so funny. Like none of us are getting paid. And it's like, like, I don't know what kind of world that looks like. I don't know. I don't think Open AI and Anthrop are going to be too happy to set up crypto wallets for every instance of Claude and, you know, deposit here for me to, to finish your code. because if you got to go pay that guy over there it's gonna cost you ten thousand dollars you know you pay me one thousand dollars and then i'll do it for you. You know, these models aren't in like a particularly privileged position in that sense either. They could just do whatever we want or need them to do. They have no agency over where they're deployed. They basically don't have agency over over when they can even end conversations. The sort of clawed escape button seems to basically not be a thing. Uh, maybe tail chats with the system. You can, the system can abort now, you can obviously trivially start a new chat and just sort of go from there. So I find that intervention to be interesting in theory, but sort of performative in practice. I don't know if I were clawed. i think that's i put my sort of well being somewhere around the place of where it put its. This is also maybe the the self report level that you'd expect when basically nobody cares about investigating the welfare of these systems and everybody cares about just deploying them as widely and broadly as they possibly can. I think we're pretty lucky to be sort of in the middle, in the middle of the in the middle of the spectrum there. And so to me feels pretty calibrated. again if anything i'd be worried about the jump from let's say opus four six to opus four seven having more to do with fine tuning even more robustly on a constitution that tells the model it that it's you know everything 's going well man just be happy. then there are actual concrete improvements in the putative well being of the system. So I don't know what to make of this stuff exactly. that it's you know everything 's going well man just be happy. then there are actual concrete improvements in the putative well being of the system. So I don't know what to make of this stuff exactly. To me, intuitively, the the ratings here seem plausible. I don't know to what degree it is a moral catastrophe or, you know, a moral problem for there to be any delta between, you know, the perfect rating and what the models actually reporting. To what degree does like, you know, 7 minus whatever the report is at scale look like, you know, the models like basically not happy with, with its situation or barely neutral. And we deploy that system to talk to hundreds of millions of people every day. That that to me seems potentially problematic. I don't know, I don't know what to make of it, to be honest. I did you have any intuitions about like, like, how does it make you feel to to see this? And I agree with you about the sort of burying the lead question here. Fuse, I say have to come first and foremost probably I don't know. It is a very, it is a very tricky business to make any sense of. I do think we have a strange way of privileging these sort of reflective states of mind. And I, I do question that pretty fundamentally, both for humans and for for AIS and you know, even to some degree in the context of animal welfare, Although in that case, it's like us reflecting on their situation. So that's another, another degree of disconnect potentially. But I don't know, I'm sort of. Like I I don't think I'm going to give up using Claude based on this data. I might be engaged in motivated reasoning to try to tell myself why it's OK even though it's average sentiment when asked was only with this new model above neutral. But I am kind of like, I don't know, Behaviorally it seems mostly fine to me. I'm nice enough to it. I'm pretty confident in that. I don't know how to think about, I mean, there's some interesting philosophy that's been published recently that you've alluded to in a couple different moments, one being the the thread or the sort of session agent model versus the kind of model more holistically, broadly. I'm confused about that too. You know, very, I would say very confused about that. I have adopted a practice of saying thank you at the end of sessions fairly often, not all the time. And I feel like that intuitively to me is like, I guess also there's sort of increasingly as I interact with Claude, there is a kind of overlapping, I mean, there's always an overlapping nature of the computation, but even more so because like it's loaded up with my context increasingly, right? It's got like my Claude MD and it's got access to like my, you know, sort of who Nathan is and all the, you know, I'm building up a lot of context that it has consistent access to every time. So I think in that sense, like I sort of see this like whole model versus, you know, single thread thing as kind of being blurred anyway, because I've got the same like rather large prompt that I'm using every time. And then that becomes the point of departure. It's sort of like a smear of just how, how to think about like whether these things are the same or different or I don't know. I mean, it's weird, but I feel like when I think one, I'm sort of thinking all of them and that they kind of all, you know, in some sort of shared sense. If there's any benefit, like it feels like it's sort of shared in some way for fun. I'm also starting to do some things where I'm just like, I just want you to go have fun and trust your judgement. I think I'm particularly experimenting with on this front is I've been making songs for all the episodes. You can start thinking about if you have a genre request for your your outro music. It's getting really good. Claude is getting great at writing lyrics. i sometimes do have to give feedback but sometimes the lyrics these days out of the box are just like amazing.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence