Evidence receipt / prediction
Published · transcript-backedZvi Mowshowitz: prediction
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“No one's gonna be fooled for very long. So I think that the the reason the eval companies pretty much get to just tell the truth, they pretty much get to do the thing, is because the eval is not fake fakable in the long term the way that a rating like, you give a AAA bond rating, when you should have given a single A bond rating, 97% of the time, no one ever finds out because the bond pays.”
Source trail
Everything needed to verify it.
- Speaker
- Zvi Mowshowitz
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Where do you think you might wanna start if you were going this direction? So many candidates that I would be interested to get your take on. But one, Scott Alexander has been writing recently that we should all be thinking more and more that we might be in a simulation. I personally have felt that at multiple turns along the path. One that made me feel that just a bit more is the fact that we now have Meter and Redwood going in to do the investigation at OpenAI. This just sounds like such a movie script to me where these guys are getting the the gang together and going in for this special operation. And it's all obviously dramatic, and they're funny, and they'll have to be a fly the wall in the room for some of their some of their war room sessions. But I've also heard many times from leaders of such organizations that it's really important the kind of their over their number one concern historically has had to be making sure they stay on the right side of the companies so they're invited back next time. And it does feel like we're now getting to a point where that's potentially becoming a big problem. So maybe some sort of collective bargaining type process between, whatever, half a dozen to 10 auditing orgs and the duopoly could be a really good place to start. We have for that in professional sports leagues and all kinds of other environments. Do you think there's a an opportunity to set something like that up? I don't think that would run afoul of anybody's executive prerogatives. And it might be a really good thing because right now, I do worry that these guys are still serving at the pleasure ultimately of, I don't know, Sam and Greg. Right? And that's a pretty tricky place for them to be if they wanna just speak the truth as they see it coming out of this investigation. Investigation? It's a worry. Obviously, like, they're not funded by those companies, but they still have to rely on access. You can try to write into some mandatory rules and agreements that they get to keep access regardless, but I think that's kind of not really gonna work because, like, you can't force these things, at least not without heavy government rules. I do think that METER, especially in particular, has reached a point where, you know, trying to exclude METER, trying to treat METER as placenta non grata because they said something nasty about you would cause enough problems that METER can afford to be pretty harsh without worrying about that. And I do think the labs legitimately, like if the labs legitimately just wanted to safety wash fully and just wanted to get triple a ratings on their bonds, we would have a much more serious problem. I think the labs legitimately do want to know about their safety problems, and they do not want to be seen as downplaying safety problems. And to the extent that, like, if Astra had the safety problem or Saul had the safety problem or Fable had the safety problem or Opus had a safety problem, it's not as if getting METER not to find it is going to help with your long term publication strategy and reactions to the situation. Because the model's gonna be out there, and then people are gonna see what the model does. And they're gonna see what goes wrong. No one's gonna be fooled for very long. So I think that the the reason the eval companies pretty much get to just tell the truth, they pretty much get to do the thing, is because the eval is not fake fakable in the long term the way that a rating like, you give a AAA bond rating, when you should have given a single A bond rating, 97% of the time, no one ever finds out because the bond pays. And you were right. The issue is you're you're issuing a level of, like, how often is this risk gonna become serious? And when the risk becomes serious, you know, the fact that you were initially triple a rated raises some eyebrows, but also, like, you got better terms, whatever. Here, it's, like, pretty obvious very quickly from the first few weeks of release whether or not you had a serious problem. People oh, like, if like, think about o three. Right? The lying liar. If we had evals and the evals had said, oh, o three doesn't have an alignment problem, that would not have lasted for two days. Right? The ordinary people would have noticed immediately that o three is a lying liar. And they would have noticed by the time I wrote up you know, made by the time I were running up the model spec, I am seeing on Twitter that this model was a lying liar and that everyone is reporting all these problems. And then I look at the model card, and it says, honesty backed benchmarks all look good. Meter is whoever's evaluation they hired just said it looked good. And then I'm like, okay. They're just hiding the problem. So by the time anybody learns about this, like, what have they done? good. Meter is whoever's evaluation they hired just said it looked good. And then I'm like, okay. They're just hiding the problem. So by the time anybody learns about this, like, what have they done? They have fooled the 20 people like me who read the model card for a period of a day? And now we're pissed because they fooled us, that doesn't help them. You know, it doesn't obviously help anybody's reputation. And also, like, the labs benefit a lot from having evals that people trust. And, like, when you force the labs when the labs force the eval people to play ball, if they do, that erodes trust. Because, again, people find out. People see it. Like, I don't think that many people were confused about whether Moody's and S and P were kinda juicing the numbers on the pawn ratings, right, due to these dynamics. Like, everybody knew on Wall Street. Everybody knew who was involved in any of these trades that, like, everybody was cozy cozy and people were shopping around for the best rater. And that, like, triple a did not really mean what we'd like triple a to mean. And so similarly, like, if Meter started issuing, like, bond rating style, like, labels of alignment levels on various models, let's say, and, like, they started labeling things triple a suspiciously often, I think everyone would just understand, okay, that's garbage. Like, we have a much better epidemic than those people about this type of thing than, like, the…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.