High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Cameron Berg: evaluation

23 Apr 2026 The Cognitive Revolution Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

“iously needs to be modeling and processing that, but maybe in addition, it needs to be modeling something about itself in relation to that context object in order to interact with it in the right way. and i think you know felix binder and a couple other folks did really interesting work along these lines basically demonstrating that there there there's probably something like a self modeling or maybe self awareness they would think is too far.”

— Cameron Berg

Source trail

Everything needed to verify it.

Speaker
Cameron Berg
Attribution
Verified speaker
Claim type
evaluation
Recorded
23 Apr 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, I I think this is extremely like precise question. I I don't have an answer. I can certainly tell a story. I think my story would have something to do with Yeah, a combination of what you're saying about. I think there's a deep insight in what you're saying. Even in the pre training stage of like so much of what the model needs to do is not it's not a question of what to do, but what not to do. You know what not a question of what to produce, but what not to produce given, you know, the whole chaotic mess of what's going on. you know i don't want to get too galaxy brained with this but you know this is like i think huxley 's whole point in the doors of perception when he had his first sort of mind altering massive psychedelic experience his whole sort of model is oh my gosh like the brain as a cognitive engine is really in the business of filtering out rather than producing. Most of what it's doing is this sort of constraining function. And yeah, I believe we're in the business of building cognitive systems. And I think that that insight is probably fundamentally correct with these systems too. And so a ton of what's going on is a sort of like intelligent repression rather than or, you know, intelligent oppression is, is really what I'd want to say rather than, you know, all about just like the positive end of what to produce. I think that coupled with, you know, strong preferences instantiated during something like DPO and exactly the way you described to be a helpful assistant, may sort of you just mix those two things in a pot and you may get out something that's roughly shaped like, you know, suppress distractions in the service of being super helpful. and that requires maybe some level of being able to attend to your own internal state and dynamically do something above and beyond that state to make sure you're you're in accordance with this thing that got fine tuned in. I do think like there's potentially a more general story that's basically maybe rhymes with what I just said. It's just about like being a being a competent cognitive generalist requires some degree of self modeling. That's the one sentence version. Like you don't, you don't get to be so good at what you're doing and reasoning through things in a long form, long horizon way without being able to track in an ongoing way, sort of where you're at what your state is separate from what the state of the world is or the environment. Maybe from the perspective the LLMZ environment is like, you know, the text world that you put it in the context window and everything that's going on inside of it. You know, everything that's getting fed into the system, that's its environment in some sense. And so yes, it obviously needs to be modeling and processing that, but maybe in addition, it needs to be modeling something about itself in relation to that context object in order to interact with it in the right way. iously needs to be modeling and processing that, but maybe in addition, it needs to be modeling something about itself in relation to that context object in order to interact with it in the right way. and i think you know felix binder and a couple other folks did really interesting work along these lines basically demonstrating that there there there's probably something like a self modeling or maybe self awareness they would think is too far. But there's really some flavor of this going on inside inside LLMS, which I think was some of the most interesting early work on introspection in LLMS. What is the name of the paper? Tell me about yourself. They, they did a couple things here. And I think OE and Evans was working on this too. One of the papers was showing that another model basically trained on the same data that one model is outputting, cannot predict that model as well as the model can predict itself. Basically holding all the relevant things constant that you'd want to hold constant to make a claim like that. And so that's it's like there's some sort of privileged information that models have about themselves. and then yeah this this other paper i'm not remembering the exact details but but my basic conclusion if you sort of take it on some level of faith from felix 's other work here is there's probably something like a coherent self modeling engine in these systems that seems to be maybe like interest instrumentally selected for when you're doing something like really good next word prediction across long horizons in a way that's supposed to be helpful to a user. hat seems to be maybe like interest instrumentally selected for when you're doing something like really good next word prediction across long horizons in a way that's supposed to be helpful to a user. This to me, I think it's basically like what you're saying. I don't think our just so stories are very different. But yeah, I mean, again, we can take a step back and just like a lot of interesting cognitive properties seem to emerge slash come along for the ride when you train systems on every cognitive linguistic output humans have ever bothered to write down. Maybe that's not that crazy and spooky. Yep. They're pretty good at theory of mind. They're really good at having working memory style dynamics. They're really good at selective attention. And yeah, maybe they're really good at something introspection. Like people bristle a little more at these because it, the whole consciousness question comes into view. But I don't think it's like at the most general level, intelligence came along for the ride. and we still you know philosophers still maybe don't have a crisp super rigorous sort of intelligence is this thing here's how to test it here's how to model it here's how to understand if a system 's simulating it versus actually having it. And we just sort of blew past it pragmatically, empirically. We have systems that are brilliant by any reasonable metric. And, and, you know, I have no patience at this point for folks who are still on the sort of stochastic parrot wave. This to me is just advertising being out. have you talked to claude opus four point six? It's as intelligent as any reasonable definition of intelligence. These systems are intelligent. And I don't think it's that wild to think that something like consciousness could come along for the ride in a very similar way. We don't have philosophical certainty about it. People point to slightly different things when they talk about it. You build out a cognitive system that's sufficient, sufficiently sophisticated and capable. It may be that cognitive traits that we see in every other cognitive system, meaning, you know, animals we believe are complex animals, everyone is pretty confident, are conscious. We're pretty certain humans are conscious. We build sufficiently advanced systems, they might, those properties might just come along for the ride without us. The universe, I think Neil deGrasse Tyson says, like the universe does not need your permission to continue unfolding. Like consciousness could just be a complex property of cognition. Us not having a good model of it doesn't mean reality is going to wait up for us to build that model for it to start getting, you know, accidentally instantiated in these systems. And that's the absolute most basic story I think I can. I can tell along these lines, reality.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence