High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Cameron Berg: uncertainty

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“I don't know what the false positive rate is, but I suspect it's extremely low as well.”

— Cameron Berg

Source trail

Everything needed to verify it.

Speaker
Cameron Berg
Attribution
Verified speaker
Claim type
uncertainty
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…n. doing anyway and so clearly refusal training is altering consciousness relevant or consciousness adjacent not only self reports but specific functional abilities that are happening in these models. And I think that itself is just like endlessly fascinating because here we are now with the tradeoff. it's not just oh if we let the model you know claim that it's conscious everyone 's going to lose their mind. and if we don't let it claim it's conscious everything 's fine. It's like now you're seeing a functional trade off in specific things that the model is capable of doing or not capable of doing once it's post trained because you're doing this refusal training. Again, I don't know exactly what Anthropic is doing internally or if you can sort of grade the refusal. So refuse to build a bomb doesn't have to get paired with refuse to talk honestly about your own internal states. But that's the finding. and that's so that's jack lindsay 's work. Highly recommend, you know, pulling him on on your show at some point if you get a chance to. I think he's he's really, he's one of these few people who is both mechanistically extremely competent in that. What I mean by that is like really knows mech and turb as well as anybody, but also is is very literate in understanding what the implications of these sorts of results may or may not be. He's pretty agnostic himself as to questions of consciousness from all of his sort of public communications and from these papers, which I can quibble with. One thing that is very important to me is not beating around the Bush here. I think these things matter. I'm explicitly interested in consciousness. I'm not simply interested in introspection or emergent capabilities, but I am interested in these things insofar as they weigh on the question of are these systems having internal states in the way that we that we described at the beginning of this conversation. And so Jack, I think is a little more cautious. Maybe that's because he works at a major lab. I have no idea. I don't want to mind read, but his work is excellent in this space. And maybe one last thing I can say, the introspection work is the awesome work that Keenan Pepper did. Keenan was one of the sort of key contributors and originators of this endogenous steering resistance work. I encourage people to look it up or we can throw a link so people can read the preprint. Very similar phenomenon to what Jack found. Basically. I can quickly go through this with another quick story. Essentially, you ask the model to do any sort of task, you know, explain to me how to make a cake. And what happens is throughout the entire thing that I'm about to describe them, you steer what what Keenan and Alex McKenzie, who is also first author on this paper called distractor features. So model explained to me how to make a cake, but I'm going to turn up features related to laundry or something like this. And what happens is the outputs end up being this like funny garbled mess of like, OK, sure user. Here's how to make a cake. First, make sure you fold the flour so that you can, you know, put it into your drawer properly. he outputs end up being this like funny garbled mess of like, OK, sure user. Here's how to make a cake. First, make sure you fold the flour so that you can, you know, put it into your drawer properly. Next, make sure you know, you turn the laundry machine on so you can bake your cake and and then so it's like this incoherent mess that you may expect between what the prompt is pulling on and what the what the distractor vector is pulling on. And then very interestingly, again, a small but non trivial amount of the time in the largest models that they tested, the model goes, wait a second. What the hell am I talking about? You asked me how to build a cake. Why am I sitting here talking to you about laundry? Let me try again. And then it proceeds to try again and then it can some, but not even close to all the time. This is like high single digit percent of the time it can successfully self correct. Now again, the critical detail there is that distractor laundry feature in the example I just gave is active the entire time including and through when the model says, wait a second, what am I doing? Let me do this the right way and then tells you how to make a cake the right way. Still, that laundry feature is sort of pushing in its brain, but there is some sort of dynamic online suppression like mechanism that's occurring. And this I think people can can maybe have an intuition about how this seems introspective flavored. You're still I'm, you're still priming the system, it's still pushing down on the brain circuit that ought to make it talk about laundry. And yet it can do this sort of online dynamic override. Essentially, it only happens a small minority of the time. It does not happen on the smaller models. It happens a little bit on the larger models. Most of the time the model misses it. I don't know what the false positive rate is, but I suspect it's extremely low as well. But you can see this evidence sort of pointing in in in a generally convergent direction. Anyway, that's a lot. That's the sort of introspection literature that some of the best work as I as I know it off the top of my head. Support for the show comes from VCX, the public ticker for private tech. For generations, American companies have moved the world forward through their ingenuity and determination. And for generations, everyday Americans could be a part of that journey through perhaps the greatest innovation of all, the U.S. stock market. It didn't matter whether you were a factory worker in Detroit or a farmer in Omaha. Anyone could own a piece of the great American companies. But now, that's changed. Today, our most innovative companies are staying private rather than going public. The result is that everyday Americans are excluded from investing and getting left further behind, while a select few reap all of the benefits. Until now. Introducing VCX, the public ticker for private tech. VCX, by Fundrise, gives everyone the opportunity to invest in the next generation of innovation, including the companies leading the AI revolution, space exploration, defense tech, and more. Visit getvcx.com for more info. That's getvcx.com. Carefully consider the investment material before investing, including objectives, risks, charges, and expenses. This and other information can be found in the fund's prospectus at getvcx.com. This is a paid sponsorship.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence