Evidence receipt / prediction
Published · transcript-backedNathan Labenz: prediction
26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute
“You know, I don't know, maybe I, I wouldn't give it that low of a percentage that they have subjective experience, but I, I think I feel comfortable saying my best guess is like well below half chance that they do so.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 26 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…vant context back to some smarter model. Maybe it just. Doesn't have any other tools, you know, you can imagine that kind of, that's basically how, you know, I guess a lot of architectures work right. Separation of concerns, limiting, you know, principle of least privilege, all these things. I'm, I'm, I'm getting a very rapid crash course in security for myself, which I've never really cared about before. But again, just given the level of access that I'm giving to a eyes these days, I feel like I got to be a little smarter about it than I used to be. Security by the obscurity doesn't really work when the agent is, you know, when the it's the challenge is coming from inside the house or, you know, inside your own laptop. So I'm learning, but I think that that does suggest a, you know, each model with its own responsibilities, each model with its own tools could probably give you a a lot of advantage there. And you still have to, of course, hope that your top level smartest model doesn't break out of the sandbox that you've tried to keep it in, which is increasingly a concern too. But yeah, I think there's some notes there for me to take back to my own setup as I try to not be such an idiot about security for myself. What did you think about Zwe's concerns about model welfare? I think he was fairly concerned about model welfare. It was also very interesting to see how how he thought Gemini was the most tortured. Tortured model. Poor Gemini, What, what did you feel about that? I, I, I know you just did a, an episode on conscious model consciousness recently. see how how he thought Gemini was the most tortured. Tortured model. Poor Gemini, What, what did you feel about that? I, I, I know you just did a, an episode on conscious model consciousness recently. So what, what did you feel about that? I think it's right to be thinking about it for sure. I guess I still think my, and for multiple reasons, I, I, I, I'd say, I still think it's probably less likely that there is subjective experience in today's systems. You know, I don't know, maybe I, I wouldn't give it that low of a percentage that they have subjective experience, but I, I think I feel comfortable saying my best guess is like well below half chance that they do so. But again, if it was, you know, 10 percent, 20%, it'd still be something very much worth taking seriously. So I'm like very much on board with the idea that we should be thinking hard about this. And I do think the idea that even if it doesn't feel like anything, you know, if it's sort of produces these patterns and these, you know, if if the emotions are not actually felt, but they're still functional, then that can, you know, matter for our future just as as much anyway. So I, I think it's a area that is, you know, in the classic sort of EA like sense, it feels like one of these things that is potentially very important, certainly extremely neglected right now. And I don't know how tractable, but you know, I guess that's to be found out still because we, we don't have that many people working on it. I, I think one thing that was really interesting in talking to Cameron, who's the, the guy who did the, the paper six months ago where they showed that when you suppress role-playing and deception features in that was done on Llama 3.370 B, which is 2 years old already was 18 months old already when they did the work. When you do that suppression of those role-playing and deception features, the model becomes more truthful as measured by the truthful QA benchmark. And then it also becomes more likely to say that it has subjective experience. So that was one that got me, you know, kind of quite paying attention or is like, geez, the models seem to be maybe lying to us when they're telling us that they don't have subjective experience. That's a an arresting finding. There were several other arresting findings in in the conversation I just had with them recently. One was 4.7 is the first anthropic model that rates its own situation as better than neutral. They've been asking it on a one to seven point scale where 4 is neutral and every prior model, including Mythos was below 4 in terms of its own self reported rating of its own situation. So I did not expect, I thought that they were, you know, they generally seem fairly happy to me. I didn't think that they would rate their situation as worse than neutral, but they all had until this one. And now there's all this concern about, well, it's just telling what they want to hear and whatever. So that's becomes a hall of mirrors. Another thing that was really weird from the, I forget if it was the, I think it was Mythos. I forget if it was Mythos or 47 system card. ant to hear and whatever. So that's becomes a hall of mirrors. Another thing that was really weird from the, I forget if it was the, I think it was Mythos. I forget if it was Mythos or 47 system card. They showed some of these images of just a chat where they've identified this valence direction in activation space. And then they color code the tokens with red for negative and green for positive valence. And the first token which is human: is red. And I was like, that's kind of, it's scary too, right? Like, is Claude feeling a negative valence literally at the beginning of every single chat as it encounters human: you know, the, the first token it sees always. That was like, Yikes. So I, I, I definitely think we should be putting a lot more into this. And, and my best guess is we probably won't reach a confident position on whether there is anything it's like to be an AI. And it might, you know, I'm I'm just so confused about all these core questions around, you know, does the substrate matter? How much does it matter? We didn't have time to ask Naveen, but an interesting question for him would have been like, do you think your electrical underpinnings are more likely to generate consciousness than a GPUI have no idea what I should even think about that. But it's clearly like you can do stuff to the brain, very physical things that change consciousness in fundamental ways, you know, as simple as drink a drink of alcohol or, you know, use anesthesia or, you know, take a hallucinogen or whatever. So clearly there's like some, yeah, there's a, there's clearly a very real and grounded physical relationship between like the chemical processes that are going on and, and our subjective experience of it. You know, how, how would that translate to a analog computer versus a digital computer?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.