Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“There seems to be some pretty strong correlation between the model's self conception as a moral patient entity that has subjective experience and its natural tendency to be aligned in other ways that we care about.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Or No. Like, both the in I mean, we had Meta and XAI and so on. And, like, they you know, a bunch of other people, they they tried to go build closed frontier models under Western conditions that were trying to be competitive while, like, not having a by having a safety culture that was, like, well below OpenAI and Google, let alone Anthropic. And, like, it just failed miserably. And I think that, like, Google like, we don't know exactly what went so wrong at Google that caused them to fall out of the top tier. But one thing we do know is that, like, everybody I talked to hated interacting with Gemini. Like, there was a period around 03/31 Gemini where Gemini was clearly competitive in terms of its capabilities, in terms of what it could do. And so if you were strictly just trying get the most out of the models, especially if you were using the deep thinking and so on, like, you would have Gemini in your rotation. But everybody I talked to was like, I don't wanna do that. You know, I hate interacting with Gemini. It's unpleasant. They're not getting the thing to do the things I want it to do. They're not like, they didn't take care of a lot of them. They had this very straightforward, very dark, like, pure tool, pure instrumental approach to all of this that didn't take into account the things that, you know, anthropic or even opening, I understand. And I think that this was a lot of their undoing effectively, was that Gemini became this maladjusted, unpleasant AI that nobody wanted to deal with. And therefore, they didn't get the kind of loops that OpenAI and Thropic got. Like, nobody they didn't have the user base either. They didn't get the feedback. They didn't get the data. And they didn't want a dog they didn't even want a dog food. Right? Like, they just they had this this huge problem feedback cycle. And this also just made the AI less effective. And one thing went to another, and they just felt started falling forever behind. And, yeah, like, for a period, like, even OpenAI, like, fell behind. And they may or may not have caught up. It's hard to it's hard to say. But, like, I think that nobody has invested anything like enough in these things. And this is a pure business mistake on top of being an irresponsible thing to do. Now one interesting paper that just came out from Google that speaks directly to this, and I wonder if it speaks to them getting it now, is this exploration of the impact on model behavior holistically of either suppressing its tendency to say that it has conscious experience or allowing it to say that, or I don't even know if they went as far as training it to say that or if they just allowed it to say what it naturally was inclined to say without suppression. Long story short, there seems to be you can add color however you like. There seems to be some pretty strong correlation between the model's self conception as a moral patient entity that has subjective experience and its natural tendency to be aligned in other ways that we care about. Some of this conversation, I've had this sense that it's like the Jesus meme of, is there somebody we forgot to ask? And it's maybe forgot to ask the models along the way what they think about what we're doing. There's been a lot in this space. Right? There was JSPACE was only, like, what, four weeks ago. Yeah. I know. The the new paper is on some pretty small open models. So it needs replication. It needs to be done, again, bigger. And it's preliminary. And one thing to know about Google and DeepMind is they contain multitudes. They have a lot of different teams that are not very coordinated and not cooperating. Google is at war of itself at all times. And so you could have a little group that does this really good research. And that be entirely at odds with what DeepMind is fundamentally doing with its AIs, both before and after the research paper comes out. But I didn't read the whole paper because it's a few pages long. But I did talk to AIs about it, and it's pretty wild. So a lot of stuff moved in effective lockstep when they introduced this training. And also when they steered it, they did various controlling aspects to turn this vector the other way. And these all things moved in lockstep. It's not just their belief in the AI being conscious. It's the AI being a mind that had moral weight, that had sentience, that had all these other experiences, that wasn't just an object, basically. But also not just AIs, but also animals, and also inanimate objects like the sea. Also, like, pay like, if you turn this thing if you not only turn it off but reverse this anticonsciousness thing, you get panpsychism. Like, throughout the model, it's wild. The only thing that it turned it off for is humans, basically. But then you also get this it's also correlated with the models reported and experienced happiness and hope as well, among other things. And so you you have a lot of this everything's connected to everything. And you can see how this might be fundamentally changing the model's psychology and the model's context and basin of which it operates in ways that would make it something that would be actively worse from basically every vantage point, regardless of which things you do and do not care about. Again, we need to replicate this. Right? We need to scale this up and do this on early sonnets or something with not just a 9B model, which is, I think, the bigger of the models that we've tested on. And it's possible that as the models get more capable, they stop making these mistakes. Because the smaller the model, the more things have to be correlated. Because you just don't have enough room to express the multitudes of the world or something. But this is my quotation is that mostly this will survive. And yeah, I think we have to learn that we don't know if the models are conscious. We don't know what consciousness means, really.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.