Evidence receipt / commitment
Published · transcript-backedZvi Mowshowitz: commitment
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“I think that people are just making a mistake. And so, you know, that's one of the reasons why I do have a lot of hope that we will do reasonable things is because, you know, I think that we're in a position we're in because only the people who invested heavily in these things were able to succeed to some extent.”
Source trail
Everything needed to verify it.
- Speaker
- Zvi Mowshowitz
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…I got some hope of informally people talk and people realize these things are bad, and they're like, we're not gonna but also, I hope they're doing that now without needing to an agreement. OpenAI clearly is doing things that are of the Buddhist of the Buddha nature of train on making the most money on the internet. Like, they're not doing that literal thing. I don't think they're that foolish. And I don't think it's that easy to do it that way. But clearly, they are making that fundamental mistake somewhere in their training process with a probability of 9.9 something. And so they need to fix that. But, yeah, I I think that they are definitely, like, trying not to do the maximally stupid things. And they're aware of the maximally stupid. But, like, again, like, it's very easy to have things that, like, you are trying not to do creep into your training process and end up happening anyway? You could say, we're not going to have training processes that have reward misaligned behaviors. How do you do that? The way that you do that is you get your act together, and you are very, very careful and prudent about all of your RL environments and all of your different training problems. And you keep an eagle's eye out for this stuff on many levels, and you make sure that it works. There's no specific thing you can say. I mean, in theory, could have a thing of like, oh, if I just we agree that if we're gonna have these various cross tests, and we're gonna, like, check each other's environments, and, like, we will throw out any environment that, like, anybody identifies problems with. And if we discover an AI has been trained on these environments, we're gonna reach we go and want rewind to a previous checkpoint before that happened and start again, so we'd better be really careful because otherwise, we're gonna waste a lot of resources. Can do these things, but, like yeah. Yeah. Look. It's it's really tough. And, obviously, one one potential advantage is if we really are in a two horse race at this point, where the two horses are pretty compatible with each other and talk to each other and they kind of understand each other pretty well, then it becomes pretty reasonable for them to make a decision that, like, not entirely opaque to each other of how much they're going to just, like, be more prudent even if this means that the action goes somewhat slower. Because, like, that also just again, my my my fundamental belief is that within a year and almost certainly probably even six months, if you invest more in fixing these problems and addressing these problems and, like, being more prudent about these problems, you will end up with a model that has better utility in the marketplace than the person who didn't do that, even if you are trading off those resources against your your capability developer. I think that people are just making a mistake. better utility in the marketplace than the person who didn't do that, even if you are trading off those resources against your your capability developer. I think that people are just making a mistake. And so, you know, that's one of the reasons why I do have a lot of hope that we will do reasonable things is because, you know, I think that we're in a position we're in because only the people who invested heavily in these things were able to succeed to some extent. You know, like, OpenAI, not it's not acting as responsibly as anthropic, and anthropic is not acting irresponsibly as I would think as the minimum level of acceptable responsibility. But I think it's pretty telling that a bunch of people tried to say this shit didn't matter, and they just had to go as fast as possible, and they all blew up. What do you mean by blew up there? You just mean, like, these incidents?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.