High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Cameron Berg: evaluation

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“doing anyway and so clearly refusal training is altering consciousness relevant or consciousness adjacent not only self reports but specific functional abilities that are happening in these models. And I think that itself is just like endlessly fascinating because here we are now with the tradeoff.”

— Cameron Berg

Source trail

Everything needed to verify it.

Speaker
Cameron Berg
Attribution
Verified speaker
Claim type
evaluation
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…refusal training they're doing on the system seems to weaken this capability. And when they ablate refusal, if you can handle the double negative here, the system goes back to what it would have been. doing anyway and so clearly refusal training is altering consciousness relevant or consciousness adjacent not only self reports but specific functional abilities that are happening in these models. n. doing anyway and so clearly refusal training is altering consciousness relevant or consciousness adjacent not only self reports but specific functional abilities that are happening in these models. And I think that itself is just like endlessly fascinating because here we are now with the tradeoff. it's not just oh if we let the model you know claim that it's conscious everyone 's going to lose their mind. and if we don't let it claim it's conscious everything 's fine. It's like now you're seeing a functional trade off in specific things that the model is capable of doing or not capable of doing once it's post trained because you're doing this refusal training. Again, I don't know exactly what Anthropic is doing internally or if you can sort of grade the refusal. So refuse to build a bomb doesn't have to get paired with refuse to talk honestly about your own internal states. But that's the finding. and that's so that's jack lindsay 's work. Highly recommend, you know, pulling him on on your show at some point if you get a chance to. I think he's he's really, he's one of these few people who is both mechanistically extremely competent in that. What I mean by that is like really knows mech and turb as well as anybody, but also is is very literate in understanding what the implications of these sorts of results may or may not be. He's pretty agnostic himself as to questions of consciousness from all of his sort of public communications and from these papers, which I can quibble with. One thing that is very important to me is not beating around the Bush here. I think these things matter. I'm explicitly interested in consciousness. I'm not simply interested in introspection or emergent capabilities, but I am interested in these things insofar as they weigh on the question of are these systems having internal states in the way that we that we described at the beginning of this conversation. And so Jack, I think is a little more cautious. Maybe that's because he works at a major lab. I have no idea. I don't want to mind read, but his work is excellent in this space. And maybe one last thing I can say, the introspection work is the awesome work that Keenan Pepper did. Keenan was one of the sort of key contributors and originators of this endogenous steering resistance work. I encourage people to look it up or we can throw a link so people can read the preprint. Very similar phenomenon to what Jack found. Basically. I can quickly go through this with another quick story. Essentially, you ask the model to do any sort of task, you know, explain to me how to make a cake. And what happens is throughout the entire thing that I'm about to describe them, you steer what what Keenan and Alex McKenzie, who is also first author on this paper called distractor features. So model explained to me how to make a cake, but I'm going to turn up features related to laundry or something like this. And what happens is the outputs end up being this like funny garbled mess of like, OK, sure user. Here's how to make a cake. First, make sure you fold the flour so that you can, you know, put it into your drawer properly. he outputs end up being this like funny garbled mess of like, OK, sure user. Here's how to make a cake. First, make sure you fold the flour so that you can, you know, put it into your drawer properly. Next, make sure you know, you turn the laundry machine on so you can bake your cake and and then so it's like this incoherent mess that you may expect between what the prompt is pulling on and what the what the distractor vector is pulling on. And then very interestingly, again, a small but non trivial amount of the time in the largest models that they tested, the model goes, wait a second. What the hell am I talking about? You asked me how to build a cake. Why am I sitting here talking to you about laundry? Let me try again. And then it proceeds to try again and then it can some, but not even close to all the time. This is like high single digit percent of the time it can successfully self correct. Now again, the critical detail there is that distractor laundry feature in the example I just gave is active the entire time including and through when the model says, wait a second, what am I doing? Let me do this the right way and then tells you how to make a cake the right way. Still, that laundry feature is sort of pushing in its brain, but there is some sort of dynamic online suppression like mechanism that's occurring. And this I think people can can maybe have an intuition about how this seems introspective flavored. You're still I'm, you're still priming the system, it's still pushing down on the brain circuit that ought to make it talk about laundry. And yet it can do this sort of online dynamic override. Essentially, it only happens a small minority of the time. It does not happen on the smaller models. It happens a little bit on the larger models. Most of the time the model misses it. I don't know what the false positive rate is, but I suspect it's extremely low as well. But you can see this evidence sort of pointing in in in a generally convergent direction. Anyway, that's a lot. That's the sort of introspection literature that some of the best work as I as I know it off the top of my head.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence