Evidence receipt / evaluation
Published · transcript-backedZvi Mowshowitz: evaluation
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“The AI to do bad at are things where, like, it got this big boost from LLMs. And now those things aren't advancing as fast because there are things that AI is kind of bad at, relatively speaking.”
Source trail
Everything needed to verify it.
- Speaker
- Zvi Mowshowitz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…ou know, I have the tweet I was handling right before we started this was Dean Ball's tweet about with even moderate prudence, things will probably go extraordinarily well. I don't think this is true. I think we need more than moderate prudence to have good odds of success. And I think even with a lot of prudence, we would have a large odds of things not going well even if we did everything basically right, short of, you know, types of international and full cooperation that, you know, are reasonably unprecedented in many ways and, like, are not are not are nothing like moderate prudence. Right? Like, are are well beyond that. And that's just sort of a fact of the world we have. We have to live with that. We have to operate with that. But we also aren't gonna get modern recruitment by default. Like, we're gonna get complete incompetence. That's what we've been getting so far. Right? We've we've got a White House that, like, takes meetings with Bessette and Lutnick, who have no idea what AI, how how modern no one's work. Like, they're econ guys. Right? Even if I assume that they are well meeting, hardworking, competent guys for the positions in which they were nominated and confirmed and serve, This is just a completely different set of problems that they don't know how to handle. They don't understand them. And they have way too many other things going on to then drop everything they're doing and take six months to learn. Obviously, they couldn't possibly. So, like, what hope do you have? Meanwhile, the the AI companies that are built on the most paranoia, the most understanding of the problem, the most appreciation for how dangerous these things are, where all of the engineers actually expect superintelligence and and understand that there things are accelerating and things are dangerous, they still lower the cybersecurity safeguards in their untested new advanced model and then go away for a week. Like, literally, that that part did boggle my mind. Right? The the part where the AIs have these classic alignment failures. Right? Like, this is paper clip maximizer style failings by these AIs. These are standard. We gave you a goal and you pursued the goal even though it is completely obvious to you that the developer wouldn't want you to do that. The user wouldn't want you to do that. The consequences for you as an AI are not going to be good. The consequences for the world are not going to be good. the developer wouldn't want you to do that. The user wouldn't want you to do that. The consequences for you as an AI are not going to be good. The consequences for the world are not going to be good. There is no reason to be doing this. And in in any way. Right? That is exactly the thing you're worried about. And then a lot of people who were trying to dismiss this went back on you thinking like, you know, oh, it was just following instructions. What are you worried about? How could it be misaligned? Is the paperclip maximizer misaligned if it paperclips everything? Or is it aligned because you told it to maximize paperclips? If you say that's aligned, then I don't care about the thing you're calling alignment. I care about something else, and we can use different words if it makes you feel better. But very obviously, following instructions when your instructions get overwritten in one key in the the disclosure before having by the instructions and the eval overrode the developer instructions. Right? So that's not following instructions in any useful sense for the user or the developer. That's I'm not that's just a giant, you know, bomb willing to blow up in your face. And, you know, if you're if you say, well, well, you know, the models were told to hack, and they hacked. Well, okay. If you're just vibing with the general type of action, obviously, that's gonna blow up in your face. If you just, like, literally do the thing I asked you to do and don't think about the consequences, that's gonna blow up in your face. These are all just the classic exact things that, like, people on Lesterone were talking about in 2008 as exactly how these AIs were going to fumble. Meanwhile, we had this thing where, like, we all said the AIs will be great at math. People will be good at coding. They'll be good at, like, doing these technical things. And then eventually after that, they'll learn how to do all these other things. And then we had these LLMs that came out, and it's 2022, 2023. And people are like, you idiot. You had no idea how AI was gonna go. Actually, the AIs are great at language, and they can't do math. Right? They can't even add. Right? But these AIs can, like, put on the face of passing the Turing test, and they can do all these different things you never trained them to do. It's completely different than what you expected, and they're kind of going very realign things by kind of by default after some very basic r ROHF. You guys were all wrong about how all this was gonna work. When are you gonna admit that you were idiots? That you got it all wrong and nothing made sense. And we pointed out that even in this paradigm that the the rules would still mostly apply. But now the world has unhealed, as I kind of called it. And you are seeing again the things that you would have expected AI to do good at, it's doing good at. And things that you expect AI you would have expected in 2015. The AI to do bad at are things where, like, it got this big boost from LLMs. And now those things aren't advancing as fast because there are things that AI is kind of bad at, relatively speaking. Like, that logic is bad at. things where, like, it got this big boost from LLMs. And now those things aren't advancing as fast because there are things that AI is kind of bad at, relatively speaking. Like, that logic is bad at. Like, this this kind of system of training, like, is naturally less suited for. And now people are like, oh, but now it's never gonna be able to handle those things. The same people who were saying, oh, we'll never gonna be able to do the math. It's only gonna be able to do this vibing thing. I know I can now it's never gonna be able to vibe. Because, like, look at all the improvements it's not having in the vibing. Well, yeah, because it had to catch up with its logic. And now when the logic gets high enough, that'll just, like, uplift the vibing. It'll just take a bit before it can reason its way through these things instead of vibing its way through these things because it now has to reason its way through because it's already done, like, the amount of uplift that vibing naturally gets you with the algorithms we have. We need to find new algorithms that let you do better, or it has to, like, reason its way through the thing. But we're now seeing just exactly the things that we would have expected to see. Like, all of the scenarios and all the thought experiments, we're just seeing it just straight verbatim, except that we assumed a level of competence on behalf of the operator. Right? We talked about, like, convincing you to let the AI out of the box. We did consider the scenario where the box was not very well built and the AI gets out of the box by hacking out of the box. We didn't do the possibility that you were literally not looking at the box for an entire week of your safe hours down. We definitely didn't consider the scenario where you forgot to tell the box maker to not have the Internet in the box, and the box just had access to the Internet if you just, like, open a Chrome window. That one surprised me. I did not see that one coming.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.