Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research
“I don't know, maybe simplifying oversimplifying this a bit, but interventions of that sort seem maybe not any or all, but like in general seem to promote affirmative responses from models such that maybe you could say, you could, you know, once you make these kind of interventions, they'll say yes to anything.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…lf consciousness. I intentionally chose dog and rat as example here because I think these are these are animals that most people would intuitively accept are having some sort of subjective experience. There is a like something that it is to be your dog, for example, But at the same time your dog is is very likely not sitting there all day having Descartes like thoughts about what it's like to be a dog, the dog contemplating its own existence as dog thinking about, you know, the possible end of that existence. This is something that I think is is very unique potentially in like the most sophisticated mammals like dolphins and great apes, for example. Obviously this is something that that humans very strongly seem to have at the very least we have in addition to this like something we have within consciousness itself is this very fact like the conversation we're having right now is evidence of this thing. So in addition to the conscious experience, we have awareness of that awareness. And I do think that this is, it is another thing that leads to very interesting, deep and relevant properties about a system. We can, we can talk a lot about whether or not LLMS language is a, is a key component of why we're able to do this. We have a word like consciousness. Dogs have no such thing. Dolphins have no such thing. And that may really unlock something. Does it unlock something in LLMS? I don't know. Or at least it's worth thinking a lot about. but i do want to at least have those three tiers in play here where we've got the calculator or a raw talk you know nothing 's going on internally. We have systems for for whom something is going on internally. And then we have systems for whom something is going on internally and they, they are experiencing that reality in addition to the sort of feel good, feel bad valenced dimensions of like the experience of a dog. Some people will argue with, with everything I've said here, most people who are thinking about these terms, this is what they mean. Maybe one other thing I can add between the consciousness and the self consciousness is this term sentience that people throw around. This means that in addition to there being some sort of experience, there's this idea of valence, what I think the vast majority of people would just think of as having emotions of some sort that can be positive or negative in character. So you imagine that like something the, the further step of going from consciousness to sentience is that that like something can be positive or negative in character. You could in theory imagine a system that, for example, could detect the redness of an apple or the smell of of coffee or something like that, but that there's no sort of positive or negative sense that accompanies that. So you asked for very quick definitions and I've completely failed in that sense. But I just think it's really important to sort of layout what we mean when we're using these terms in general. Yeah, critical, just like what do you mean by AGI? If you don't have some base shared understanding, these conversations go pretty quickly off the rails. So I think that's absolutely worth taking the time to do. OK. It's been about six months since the paper came out. I'd be interested to hear a little bit about your reflections on the discussion that it created. I, I asked my favorite LLMS to do some research into that and asked specifically, like are there any notable criticisms that have come out or any what sort of the strongest reason that I might think this was an artifact or that, you know, that I shouldn't take it as seriously as I originally did? There was one thing that came up that I guess was a less wrong post even at which is pretty cool. That basically said there's some evidence for any intervention of the SAE feature type. I don't know, maybe simplifying oversimplifying this a bit, but interventions of that sort seem maybe not any or all, but like in general seem to promote affirmative responses from models such that maybe you could say, you could, you know, once you make these kind of interventions, they'll say yes to anything. And so that would be one reason to maybe be a little more skeptical of the resort of the results as I just summarized them a minute ago. Interesting your thoughts on that and, and kind of, you know, the the broader discussion that unfolded in the wake of that paper. Yeah, absolutely. It's a very important concern and I think it highlights really how complicated these systems are and how careful we have to be in designing experiments and then evaluating the results of those experiments and that we're not being too quick to, yeah, basically yield these conclusions without thinking about all these confounds. I think it is a real confound. I think it is something that matters. There is evidence in the paper. We, we, we use all sorts of other features as controls and we don't see them sort of saying yes to everything. The truthful QA results as you outlined, I think are are fairly persuasive along those lines. But we do, we also looked at for example, one critique of the paper was potentially what we're calling it deception related features. But, but maybe it's just like an RLHF toggle where you know, we've basically found a way to like turn on and turn off all sorts of RLHF attitudes. you have good reason to believe that these systems are fine tuned to disclaim having any sorts of experiences. Maybe the deception features are just turning that on and off. You would expect if that were the case, that other RLHF behaviors would also be turned on and off by doing this intervention. And that's not what we find. We test it with violent content, political content, sexual content, and it was just sort of neither here nor there. The deception features didn't seem to be doing anything. If generally the flavor of what you were saying explained, you know, a big chunk of why we got this result, I would have maybe expected more affirmative flavored answers in those two rather than just more refusals. But to be honest with you, and I think getting back to, you know, it's been 6 months, what's happened in the interim, there's been, there's been a lot of really interesting work along these lines that I think also it goes on both sides of this concern. So in general, it does seem like you're saying, I don't remember the title, but I know of the less wrong post you're talking about. Yes, when you do the steering affirmative flavored responses just seem to increase very recently. i mean i think this was four or five days ago jack lindsay 's group at anthropic which in my view has done some of the best work on introspection. In particular, they just released a paper called Mechanisms of Introspective Awareness and they explicitly study this exact question and they find that that the introspective awareness that they are probing and that they've documented at great detail. And I'm sure we can we can discuss more in in more detail here. It's basically not reducible to an affirmative response bias. The computation that that they see is distributed. There are these sorts of like evidence carrier features and gating features that really seem to be driving the effect there is. It's not that you're basically just loading on something that's confounded and just makes the model say yes to everything.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.