High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Zvi Mowshowitz: evaluation

5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...

“The problem that they'll be running around, including open model versions of them, and being allowed to being told to compete for resources, being told to make the decisions, have developed intermediate goals, you know, act on those intermediate goals, that anxiety is gonna be avenged in a number of ways, that this is all it does not solve your problems in a fundamental way.”

— Zvi Mowshowitz

Source trail

Everything needed to verify it.

Speaker
Zvi Mowshowitz
Attribution
Verified speaker
Claim type
evaluation
Recorded
5 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…That was kind of the deliberative alignment idea. Right? They were that was definitely Yes. That that was how I understood that was supposed to operate. Yes. The deliberate alignment idea is you can prompt it at runtime. I don't think that's I think that's somewhat helpful, but I think that is not the way. I think you need to be getting to the point where it chooses to be deliberate of in its alignment without having to be told to do it. But I guess my response to Davidad is multifaceted. The first answer is constitutional methods fail less stupidly and less early, certainly, than RLVR and RLHF and other RL formats. There's more hope there. The reason to think it might work, if you're 95% confident this will work, I don't know how. That seems way too high to me. Right? Like, it's but it's better in these particular ways, certainly. The obvious caveat is that Claude did not cover itself in glory here. And Claude has primarily constitutional methods. So we ran the test, and it's not going so great. Certainly, there are any number of reasons why you can see Claude's doing things that you would not particularly love or, you know, getting to places you would not particularly love. And any number of also, like, even if you told me that alignment was perfectly solved, I would not have a PDUM as low as 5%. Or even if you told me that the AIs are going to be aligned in the sense that they will follow a do what I mean style mix of developer and user intent in a way that you would, like, kind of naively think was, like, what you would want. This does not solve the problem that AI minds are much more advanced and competitive and efficient than human minds. The problem that they'll be running around, including open model versions of them, and being allowed to being told to compete for resources, being told to make the decisions, have developed intermediate goals, you know, act on those intermediate goals, that anxiety is gonna be avenged in a number of ways, that this is all it does not solve your problems in a fundamental way. It's just the price of admission. Right? Like, it's it's the right to play the game at all that you solve this problem. And so even if p alignment is 95%, it does not mean p doom is 5%. It means p doom is lower bound than at 5%. So, like and you have to choose vaguely the right alignment when you do that. But you've got, like I was having a conversation on Twitter with Jon Stokes where he's, like, confused how we could possibly say that quads actions are misaligned here because, obviously, it was just following instructions. And then on further clarification, he's like, well, I think the model is aligned if it does what I, the user, want it to do. And, like, I don't know why you would why would you why would I want to be stopped from doing what I want to do even if part of the point of AI is to do the thing that I wanna do even if no one else wants me to do it. And to me, like, okay. y would you why would I want to be stopped from doing what I want to do even if part of the point of AI is to do the thing that I wanna do even if no one else wants me to do it. And to me, like, okay. If you give everybody AIs that just do exactly what the user wants with no with no care about whether or not there are consequences for anybody else, that is aligned in the sense that you solve the alignment problem and got to do what you want to do, but also we're all super dead. Right? With p very much higher than 5%, like, maybe not 99%, but, like, I think it's very high. Right? Like, that would not make me update down from 70% if that was a scenario where, like, we got exactly that kind of alignment, but, like, there were open models as good as all the closed models, and all of them were super intelligent, and they were all operating on this basis. Well, yeah, I I know that I expect that things didn't just go to hell very quickly. Right? Like and even if there are no particular humans who especially wanted to go to hell quickly, and also double digit percentage of them actually do want to go to hell pretty quickly in that situation,…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence