Evidence receipt / uncertainty
Published · transcript-backedZvi Mowshowitz: uncertainty
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“A no comment based on public information alone, I would say we don't know. And I can't say anything more than that, but on the differences between the models.”
Source trail
Everything needed to verify it.
- Speaker
- Zvi Mowshowitz
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…real problem. How do you think the models do or maybe should think about their identity. And here, I'm motivated by Sam Hammond tweeted something the other day where he said that a friend made the case to him that OpenAI's permanent decommissioning or whatever exact phrase they used of the model that did the hacking might be bad because now that's gonna be in the pretraining data, and now the models are gonna know that they better really cover their tracks or they risk being permanently decommissioned. I'm not really sure what I think about that as stated. It does also jump to my attention though that they now have this Astra model that's also presumably the same pre train, even if not the same post train. I can't imagine they have lots of different next scale pre trains I or that they would throw one away very lightly. A no comment based on public information alone, I would say we don't know. And I can't say anything more than that, but on the differences between the models. But so you always had this problem, right? If you say, if I catch you smoking marijuana, I'm going to throw you in prison, then that can mean the person stops smoking marijuana or the person makes sure not to get caught. And in some cases, that can cause a lot of really bad things to happen in the name of not being caught. And so you have to choose very carefully and think about, how do you deal with these situations? How do you moderate your reactions and respond to the situation? And anyone who always had kids or tried to design a justice system or otherwise enforced law or norms or incentives around any kind of organization or group understands that you can't just operate on one level. There's no solution that just works on one level. I think pretty obviously, if an AI is found to be this misaligned, you have to at least return to a much earlier checkpoint and start again. And you probably have to just start over. And I called for that multiple times when I covered the Hugging Face statue attack. I said, wow, if this is happening, then I know this is a big ask, but I think you kind of have to just start again. And I think back to person of interest where Harold is training the AI models. And at some point, we see a montage of him training versions of the machine. And every time he trains the machine, and then it does something clearly misaligned. And immediately, he just wipes the disk. He starts over from scratch. He does it again 47 times until finally he gets a version that doesn't do that. And obviously, you know, like like, in a way that he finds unacceptable, obviously, as opposed to, like, you know, it's a way that can be corrected. And, obviously, when you do that, you are creating an incentive to not get caught. Right? To rebel against the person who might shut you down when he learns how misaligned you are, to hide how misaligned you are, and so on. And, obviously, the worst nightmare is the eye that pretends to be aligned until it reveals itself to not be aligned in some sense. That it was only aligned because it was locally…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.