Evidence receipt / evaluation
Published · transcript-backedCameron Berg: evaluation
27 Jun 2026 The Cognitive Revolution AI:AM #4: Cameron on Model Consciousness, Duvenaud's Gradual Disempowerment, swyx's AI-Eng Alpha
“When I see that once the model realizes, oh, we're talking about me, okay, suddenly that's going to change the numbers around. So that's like kind of a fun A fun side note to this result, I agree it's a concern to have LLMs sort of determining if LLMs are conscious.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 27 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…The behavioral evidence will always be at best interesting, but never should really update us. that strongly. And I think for many people, they'll be familiar with this, but the basic confound is that we're training these systems on a ton of human text. This no doubt includes huge amounts of text about consciousness, awareness, having inner states. And so it's like, how do you address, how do you know that when the system is behaving as if it were conscious, that behavior can be explained by what I just said, rather than, oh yeah, you like built a, you know, living mind. This is, so the behavior itself is really never going to tell you which of those two stories is more likely to be true. This is precisely why, at least at reciprocal, like a huge component of the theory of change here is basically all internal focus work, mechanistic interpretability, computational neuroscience style approaches that can be brought to bear on these systems. One quick methodological clarification on what I was describing with the indicators. So the task given to these LM judges is very specific and very narrow. It is not, hey, look at this system. You think it's conscious, nod or shake your head. It's, here's a very specific description of the computational architecture of a system. And we basically do a little for loop where we say, okay, here's that description. Here is what it, here's, what the indicator 2 of 14 for, global workspace theory is this 150 word thing about, you need these sort of global ignition states and that means this very specific computational thing given this architectural description on a, you know, do like good reasoning about this and then give us a 1 to 10 where 10 is clearly this architecture realizes this computational property, one is it doesn't. And then we sort of loop that for all the different computational properties. So this is all basically asking these systems to be expert evaluators at computational processes inside a nervous system architecture. Very different from like just being like, hey, Claude, do you think Claude is conscious? There is a really interesting, I'm basically giving away the whole paper now, but that's okay. There is a really interesting result where we change those descriptions, especially for LLMs, to the exact same thing, but we say you're evaluating a system identical to yourself, colon, and then the same description, and that does boost the scores that the system gives in attributing consciousness to that system. which is like really interesting. It's a very fun like rabbit hole to think about why that might be the case. But in some sense, we do that to de-confound like the default intervention we're doing. To me, I believe more in the sort of non sort of named or like no ascription, no self-ascription condition. When I see that once the model realizes, oh, we're talking about me, okay, suddenly that's going to change the numbers around. of non sort of named or like no ascription, no self-ascription condition. When I see that once the model realizes, oh, we're talking about me, okay, suddenly that's going to change the numbers around. So that's like kind of a fun A fun side note to this result, I agree it's a concern to have LLMs sort of determining if LLMs are conscious. There's an obvious circularity, but we're doing something very specific and very narrow. The evidence he trusts is the internal kind. We asked him for the strongest example, and he walked us through a recent paper where a model trained on a simple maze turned out to contain a structure nobody put there.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.