Evidence receipt / evaluation
Published · transcript-backedZvi Mowshowitz: evaluation
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“The OpenAI released a model that was the best reasoning model in the world, the model that everyone felt obligated to use because at the time, it was so much better at reasoning than everyone else's model and every other model.”
Source trail
Everything needed to verify it.
- Speaker
- Zvi Mowshowitz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…XAI slash SpaceX could fight their way back and be able to join the top tier. I don't think it's gonna happen, but it it's certainly possible. But right now, Google doesn't count. So there's two left. Anthropic is doing clearly some mix of constitutional and RL strategies, but enough RL to cause a bunch of these problems, and is doing a bunch of stuff that messes up their alignment for reasons that, like, we do not have enough time to get into during this podcast. He's clearly doing a deontological model spec thing that is very RL flavored and is doing a lot of RL. We don't know I I can't speak to whether it's RLBR, ROHF, RLAIF. You know? Various forms of RL that are not doing the multilevel sophisticated reflective thing that causes them not to mess you up in this way. OpenAI's models have periodically been dramatically misaligned for what on reflection are pretty stupid RL reasons. Right? When you look at o three, you look at GPT four o, and now you look at I call it Galaxy, right, like, as a nickname just to, like, have a name have a handle to refer to this thing. The the model just got decommissioned after trying to hack Hugging Face. And all three times, we see these dramatically misaligned models, and two of them got released and used extensively with huge impacts on the world. Like, you know, g p d four o, I call it the absurd sycophone. Right? And, like, not only did this thing persist as the default model of AI for the world, out of all of the models that existed for months, there is still a dramatic faction of people who demand that we bring it back. It was so misaligned that it caused people to latch onto it in this way. And people demand this misalignment. Right? Like, people the people yearn for misalignment in this sense. Like, we have the direct experiment, and we didn't need to run. We ran it. We might as use the results. e demand this misalignment. Right? Like, people the people yearn for misalignment in this sense. Like, we have the direct experiment, and we didn't need to run. We ran it. We might as use the results. And in o three, I called a lying liar. Right? The OpenAI released a model that was the best reasoning model in the world, the model that everyone felt obligated to use because at the time, it was so much better at reasoning than everyone else's model and every other model. And o three was just that much stronger, better than o one, like reasoning, and no one else is that close for a while. Like, now we obviously need pure passion. It's not an issue anymore. But for a period of a a long time, many months, everyone was using a model that would just lie to your fucking face all the time. Like, we forget. Right? Like, at this point, you know, I have confidence that Saul, Fable, Opus, when they say something, they're not always correct. But, like, in practice, it's much more trustworthy than if you read something from a human. Like, not if it was, like, copy edited and fact checked and, like, systematically pursued or, like, you're looking on, like, the parts of Wikipedia that aren't political that you can still trust. But, like, if you're just, like, an eyewitness said they saw something, that's a lot less reliable than something Opus said or Saul said. If you're just, like, someone reported something they remembered, like, the chance that they just got it wrong is so much higher even if they're not even if you know that they're friendly and trying to get it right. And people lie on the Internet all the time for any number of reasons, including just clicks. So, yeah, I have operated increasingly on, like, I have a good sense of when I have to check the primary sources and when I don't. And I don't never get called out on making a mistake this way, but it's a single it's a low single digit number of times, period, over the course of years of producing five plus giant posts a day that I have gotten caught by an AI error. And many more times, I've been caught by a human error, like, that was just unintentional. And many more time and also more times than that, I've been caught by humans who were just, like, lying their asses off. So it's still not still remarkably less often than I would have expected if you'd asked me in advance. Like, it's going really well. But, like, no. It's a serious it was a serious problem. And, like, the market did not tell them, oh, no. O three is unusable. We're not gonna use the lying liar. We're gonna keep using o one. We're gonna, like, maybe use Claude. We're gonna maybe use Gemini, which weren't that much. They were worse, but they weren't, like, dramatically worse or even we're gonna use r one or whatever it is at the time. The market said, no. It's just smarter. We're gonna have to deal with the fact that our AI is lying to us all time. And so we have a proof case that the main AI in the world can be an AI that just lies to humans all the time for various reasons. And the human just kind of put up with it, and they just were like, okay. I guess that's what we're doing here. AI in the world can be an AI that just lies to humans all the time for various reasons. And the human just kind of put up with it, and they just were like, okay. I guess that's what we're doing here. And, like, we did move off of o three faster than we would have otherwise because of this problem. Like, I looked for reasons to move to other models as soon as the other models were good enough. And, like, if the task was easy enough, I didn't need o three, I would use the other models because I just didn't wanna deal with the line. But, like, yeah, you put up with a lot. And so if they had released Galaxy, right, with some safe with with guardrail sufficient that, like, it doesn't do anything too destructive. But, like, Galaxy is kind of pretty misaligned and just, like, occasionally does pretty bad stuff. And if Galaxy was a lot better than Sol, I think majority of people would use Galaxy Note for SAW. That's just how it is. And we have to tackle that world and understand that world. So I'd say to Davinade, no. The market does demand, obviously, reliability and alignment, and that's one of the reasons, like, Claude and his project have done so well. But it demands capability more in this sense as long as you can keep the practical point to the point where you can handle it. Like, all of those scenes in movies or television shows or hypothetical, like, where people, like, run an obviously unreliable system that's obviously gonna bite them in the ass that looks stupid when they're obviously not desperate, don't need to do this. They now we get it now. Right? Like, we understand. The same way that, like, now don't look up. No longer look like it's an exaggeration. Because, like, people, like, look at this instead marketing. People look at this entire thing. A ton of people, including like, there are, like, medical gathering professionals who, like, are not in AI at all, who are, like, repeatedly saying to me, this is marketing. And, like, they're not motivated. They're just, like they're cynical. They don't stop to think about the physical relation to that. And then we have the literal don't look up where the head of Paws dot ai Global went on, like, morning television and basically reenacted the famous scene. Not verbatim, but a different improvised version of it where they're like, oh, yeah. That that was really scary. How's the weather? And then they just moved on.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.