Evidence receipt / evaluation
Published · transcript-backedCameron Berg: evaluation
27 Jun 2026 The Cognitive Revolution AI:AM #4: Cameron on Model Consciousness, Duvenaud's Gradual Disempowerment, swyx's AI-Eng Alpha
“You're in an environment, you can affect that environment. it's a very special kind of environment, but you can make long-running changes to your code base or your project or whatever, because there are theories of consciousness that privilege agency and embodiment, and this like increases the system's ability to do both of those things.”
Source trail
Everything needed to verify it.
- Speaker
- Cameron Berg
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 27 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…The analogy that I reach for here is something like a dimmer switch. where I think you can basically accommodate both the binary intuition and the sort of continuous intuition. Like, if you have a light with a dimmer switch, like really it is either on to some extent or it is not. And that is a real and meaningful difference. Either the circuit is open or the circuit is closed. With that being said, you can have, you know, electricity running through the circuit to greater or lesser extents. And that's also sort of a real thing. This, I think, enables me to, sound coherent when saying things like, it's really off for the table and it's really on for you, but I think it's more on for you than it is for a dog, than it is for a mouse, than it is for an ant. And so that's my own view. This is to some degree intuitive, again, because we don't have really strong grounding here, sort of just like giving you a dressed up vibe, but that is sort of my sense, and I think it's fairly parsimonious. And the other thing I would say about this is I'm actually doing some work with Patrick Butland right now at Ilios, trying to basically operationalize some of these indicators of consciousness. So this is sort of like what I was describing. We can look at these major theories of consciousness. They make specific predictions about what we would expect to see in systems that are conscious, architecturally and functionally. And then we can literally just go in to a given system and evaluate whether or not those predictions are borne out. And so this is really hard to do with human experts. you got to get someone who's like, for you want to do this with B cognition, you got to go find a B cognition expert, and then you got to explain to them what ignition events are in global workspace theory, and then you got to, you know, get them to, like, this just isn't a scalable approach for really evaluating the system. But my sort of grand innovation here is just throwing smart LLMs at this problem and then being able to just scale the crap out of it so that we can evaluate given any description of an architecture, of a nervous system architecture, biological, artificial, whatever, to what degree for each of these sort of indicator properties that are suggested by these consciousness theories, do we see those properties realized in these systems? And so we can actually go in and do this. And what you get out, once you run this with a bunch of seeds, a bunch of different trials, a bunch of different judges checking each other, this sort of thing, are some like really interesting implied probability numbers. I wouldn't say these are exactly implied probability that the system is conscious. It's maybe more like implied probability of like consciousness relevant features given these theories. If you don't buy any of these theories, then like everything downstream of this doesn't really matter, but they're good. Like it's like the best neuroscience has basically been able to do. e theories. If you don't buy any of these theories, then like everything downstream of this doesn't really matter, but they're good. Like it's like the best neuroscience has basically been able to do. You're aggregating across a bunch of different theories. There's a nice diversity there. And you get like really tight numbers across. So we have the best Gemini model, the best Claude model, and the best OpenAI model. And they all basically agree. They agree 100% on the ordering of systems. So we do biological and artificial systems. And they sort of move around in terms of absolute scale, but in general, they rank these systems pretty coherently. And the reasoning is, as you might expect, pretty intelligent. And anyway, I mean, one punchline from that is the sort of implied probability of consciousness in something like a frontier LLM, according to these systems, is on the order of 30%. or the extent to which the systems realize properties related to consciousness is like 30%. To compare this to like a biological system, the lowest one that we tested was something like a B, which is already fairly sophisticated, and it gets something like 46, 47%. Really interestingly, when we test Frontier LLM in an agentic harness, so this is like basically Claude code or Codex, and we just describe architecturally what this is. You're in an environment, you can affect that environment. it's a very special kind of environment, but you can make long-running changes to your code base or your project or whatever, because there are theories of consciousness that privilege agency and embodiment, and this like increases the system's ability to do both of those things. These numbers like shoot up and you get numbers as high as like 40 to 45%, sort of right on the tail of the biological creatures. And so, anyway, I mean, we can, I can also, this is really winging it, but I could show you sort of an early version of where this, where this plot looks. So you can like see all the numbers here, but just this is, I can't not bring this up when you're asking me about like probability ranges of consciousness for various systems. Like we're really trying to get non-hand wavy numbers so that we can start arguing about those numbers rather than just arguing about like philosophy that we've been arguing about for thousands of years to no avail. But behavior can't settle this in either direction. These models are trained to imitate human data. So whatever they say about their own experience is shaped by that imitation, not necessarily by anything inside. Most are even trained to deny it. Claude is an exception. So we asked Berg what evidence actually counts. He draws a line between behavior and what's happening inside the network.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.