High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Zvi Mowshowitz: evaluation

5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...

“I think we need more than moderate prudence to have good odds of success. And I think even with a lot of prudence, we would have a large odds of things not going well even if we did everything basically right, short of, you know, types of international and full cooperation that, you know, are reasonably unprecedented in many ways and, like, are not are not are nothing like moderate prudence.”

— Zvi Mowshowitz

Source trail

Everything needed to verify it.

Speaker
Zvi Mowshowitz
Attribution
Verified speaker
Claim type
evaluation
Recorded
5 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…You'd simultaneously be horrified by and grateful for these forms of complete recklessness and incompetence on the infrastructure and supervision sides by these companies. On the one hand, this is a horrible situation they absolutely have to fix. And, like, we're so fucked if we don't fix it. But on other hand, that can be fixed. And by not fixing it, we get to see these things while they're relatively harmless, while they are relatively preventable, while they are in their easy platonic forms and, like, can be appreciated. On the flip side, that that gives people the excuse of, oh, these people are just incompetent. And that can cause people to dismiss the underlying situation. So it it does work both ways. To me, like, you know, we we we've seen failure on every level. It's like, I called it a total less wrong victory in the sense that everything is going the way you predicted and a total less wrong defeat in the sense that everything is going the way you predicted. Where, like, we did predict. You know, the Elias has the law of earlier failure, which is that, you know, plan will fail at a much earlier point for much stupider and more preventable reasons than you thought it would fail even if you thought the plan would definitely fail and had good reasons why it would definitely fail. But you can't if you would explain to people two years ago, OpenAI's models are gonna be misaligned, and they're gonna go out there and they're gonna hack major websites because OpenAI will just not care if their sandboxes are misconfig are not con are not strong enough to hold the AI. The AI will break out repeatedly. They will notice this. They'll be worried about this, but they will just leave the sandbox there for the AI to break out of while the safeguards are down and they just don't look at it for an entire week. People would say, that's stupid. Nobody is that incompetent. That would never happen. And your scenario makes no sense, and they would use this to then dismiss these stupid doomer concerns or whatever. Whatever. Because, like, obviously, like, people will just but people won't just. Right? People will never be just in this sense. Right? You know, people have never just anything, and they're not gonna start now. And we need these displays of utter incompetence and dirtiness in a general sense. And, like, one of things I've been hammering hammering is if your plan cannot survive the real world level of dirtiness and incompetence and ordinary human error, then your plan is insufficiently foolproof because of all the fools, and it will definitely fail. Even if your plan would have succeeded if we were not fools, we were competent, and we were responsible. So my position basically is, you know, I have the tweet I was handling right before we started this was Dean Ball's tweet about with even moderate prudence, things will probably go extraordinarily well. I don't think this is true. ou know, I have the tweet I was handling right before we started this was Dean Ball's tweet about with even moderate prudence, things will probably go extraordinarily well. I don't think this is true. I think we need more than moderate prudence to have good odds of success. And I think even with a lot of prudence, we would have a large odds of things not going well even if we did everything basically right, short of, you know, types of international and full cooperation that, you know, are reasonably unprecedented in many ways and, like, are not are not are nothing like moderate prudence. Right? Like, are are well beyond that. And that's just sort of a fact of the world we have. We have to live with that. We have to operate with that. But we also aren't gonna get modern recruitment by default. Like, we're gonna get complete incompetence. That's what we've been getting so far. Right? We've we've got a White House that, like, takes meetings with Bessette and Lutnick, who have no idea what AI, how how modern no one's work. Like, they're econ guys. Right? Even if I assume that they are well meeting, hardworking, competent guys for the positions in which they were nominated and confirmed and serve, This is just a completely different set of problems that they don't know how to handle. They don't understand them. And they have way too many other things going on to then drop everything they're doing and take six months to learn. Obviously, they couldn't possibly. So, like, what hope do you have? Meanwhile, the the AI companies that are built on the most paranoia, the most understanding of the problem, the most appreciation for how dangerous these things are, where all of the engineers actually expect superintelligence and and understand that there things are accelerating and things are dangerous, they still lower the cybersecurity safeguards in their untested new advanced model and then go away for a week. Like, literally, that that part did boggle my mind. Right? The the part where the AIs have these classic alignment failures. Right? Like, this is paper clip maximizer style failings by these AIs. These are standard. We gave you a goal and you pursued the goal even though it is completely obvious to you that the developer wouldn't want you to do that. The user wouldn't want you to do that. The consequences for you as an AI are not going to be good. The consequences for the world are not going to be good. the developer wouldn't want you to do that. The user wouldn't want you to do that. The consequences for you as an AI are not going to be good. The consequences for the world are not going to be good. There is no reason to be doing this. And in in any way. Right? That is exactly the thing you're worried about. And then a lot of people who were trying to dismiss this went back on you thinking like, you know, oh, it was just following instructions. What are you worried about? How could it be misaligned? Is the paperclip maximizer misaligned if it paperclips everything? Or is it aligned because you told it to maximize paperclips? If you say that's aligned, then I don't care about the thing you're calling alignment. I care about something else, and we can use different words if it makes you feel better. But very obviously, following instructions when your instructions get overwritten in one key in the the disclosure before having by the instructions and the eval overrode the developer instructions. Right? So that's not following instructions in any useful sense for the user or the developer. That's I'm not that's just a giant, you know, bomb willing to blow up in your face. And, you know, if you're if you say, well, well, you know, the models were told to hack, and they hacked. Well, okay. If you're just vibing with the general type of action, obviously, that's gonna blow up in your face. If you just, like, literally do the thing I asked you to do and don't think about the consequences, that's gonna blow up in your face. These are all just the classic exact things that, like, people on Lesterone were talking about in 2008 as exactly how these AIs were going to fumble. Meanwhile, we had this thing where, like, we all said the AIs will be great at math. People will be good at coding. They'll be good at, like, doing these technical things. And then eventually after that, they'll learn how to do all these other things. And then we had these LLMs that came out, and it's 2022, 2023. And people are like, you idiot. You had no idea how AI was gonna go. Actually, the AIs are great at language, and they can't do math. Right? They can't even add. Right? But these AIs can, like, put on the face of passing the Turing test, and they can do all these different things you never trained them to do. It's completely different than what you expected, and they're kind of going very realign things by kind of by default after some very basic r ROHF. You guys were all wrong about how all this was gonna work. When are you gonna admit that you were idiots? That you got it all wrong and nothing made sense. And we pointed out that even in this paradigm that the the rules would still mostly apply. But now the world has unhealed, as I kind of called it. And you are seeing again the things that you would have expected AI to do good at, it's doing good at. And things that you expect AI you would have expected in 2015. The AI to do bad at are things where, like, it got this big boost from LLMs. And now those things aren't advancing as fast because there are things that AI is kind of bad at, relatively speaking. Like, that logic is bad at.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence