Evidence receipt / evaluation
Published · transcript-backedZvi Mowshowitz: evaluation
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“Are you gonna fix it? And so, like, this idea of, like, locking into requiring certain training techniques, like, I think there are legitimate complaints that would be, like, potentially, like, exactly the wrong thing to do and could hardly backfire because, like, the government moves so slowly and, like, you can't, like, undo those kind of requirements.”
Source trail
Everything needed to verify it.
- Speaker
- Zvi Mowshowitz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…So on these training the question of training methods, it seems like you're basically saying we don't sorry, Davidev. We don't have such a silver bullet, but we clearly do have some things that we've are seeing in pretty vivid ways are driving real problems, especially if they're driven to the extreme. And your position is like, no. The market doesn't really punish this too much. It actually just tolerates all kinds of weirdness if it's part of the package that gives you the highest end capabilities. So what do you think we might ought to do, and how might we ought to construct agreements perhaps? We have this, obviously, this letter. There seems to be some op opportunity or some appetite for the coordinated pacing of the frontier. What do you think can we come up with simple rules that people could agree to where we could another mental model I have been we're just teasing around lately is I do believe some AI risk is probably irreducible. But then also, there's a lot that we're just asking for right now. Can we come up with simple rules that at least allow us to take most of the risk that we are currently asking for off the board? Those might be things like have some ratio of flops that is, like, the max ratio you can put into RLVR versus constitutional. Or everybody has to spend so many flops doing pretraining data filtering to try to just get certain bad notions out of the mix in the first place. Are you obviously, those could backfire if they're bad ideas or send us down wrong paths. But it seems like we might wanna try something in that department. Do you have any hope for that? And what if you do, what would you propose we agree to do or not do first? So I don't think you can take, like, most of the risk off the table with those kind of strategies even if you implemented, like, wise versions of those strategies with universal agreement or anything like that. But I also don't think that you can implement that. I so, like, one of the lessons that we've had over the course of years is that there is tremendous resistance to anything but the most simple interventions and the most simple rules. Also, most simple rules and the most simple interventions, but less so. But, like, if you talk to people and, like, you try to specify certain training techniques, remember all those people who threw a fit about how we were, like, locking in a potentially inferior, like, charging regime in the in for the EU because the EU was requiring Apple to switch to USB C? Even though USB C was, like, obviously the best answer right now. And, like, Apple was just kind of being an asshole by insisting on their own plugs. Because, like, what happens if they develop a better plug? Like, what happens if it turns out USB C is is not a perfect technology? Are you gonna fix it? And so, like, this idea of, like, locking into requiring certain training techniques, like, I think there are legitimate complaints that would be, like, potentially, like, exactly the wrong thing to do and could hardly backfire because, like, the government moves so slowly and, like, you can't, like, undo those kind of requirements. But, like, you definitely can't do that. People would throw a fit about, like, trying to dictate training techniques and, like, how are you gonna enforce that? Are you gonna imagine the 10:47 debate except, like, the opposition tuned up by two orders of magnitude or something crazy. Just, like, it would be completely outside Overton window to even, like, try to make something like that stick. You can, of course, you know, try to strongly encourage doing more intelligent forms of all of this and trying to encourage people to move to different bases. Like, I've been trying to, like, not so subtly encourage OpenAI to, like, move to a constitutional virtue ethics style basis for a while now. And I got an absolute, you know, traction that I noticed. Like, who knows what they're doing internally? But, like, it doesn't seem like, if anything, they're doubling down on RL. They're doubling down on these types of methods, and, like, that's why you see what you're seeing, I would assume. I don't know what else they're doing. But they probably have some new methods we don't know about. Right? Like, the trade secrets they're not describing. But by default, basically, technique is going to probably look like the kind of bad thing that causes more problems. But if you ban a specific technique, what they come up with? Or, like, find a way around the rule is just gonna be worse. You understand? bly look like the kind of bad thing that causes more problems. But if you ban a specific technique, what they come up with? Or, like, find a way around the rule is just gonna be worse. You understand? Because, like, at least with the current techniques, we've had some years to figure out the worst possible ways to do them and to do them slightly less stupidly. Like, when you you really when you're training a mind, you have to be thinking on, like, every possible meta level at once about where the incentives are and what you're steering towards, and you have to generate a world in which you and the mind together are in some sense cooperating to identify ways in which there are feedback loops that are going in bad directions, in which you are creating bad scenarios on any level, and, like, treat various cancers kind of as inevitable thing that you have to, like, notice and stomp on. And, like, otherwise figure out how to handle all these things and create an antifragile system. And none of these things are gonna happen unless you deliberately, like, set out to do them. And I want to devote a lot of effort to doing them because you really value what you get out of that, and that will pay dividends. Like, I think that Anthropic has won tremendously to making these investments and to making 10 times as much investment as they have. But, you know, it's very hard to convince somebody to go ahead and do that. What you can do is you can set incentives to be like, no. Seriously, you're not gonna screw this up. We're gonna punish you a lot for screwing this up. Because a lot of the problem here is that, like, you just you don't pay the externality. When OpenAI wipes someone's hard drive when it's all wipes someone's hard drive, and early on, there was a problem with Salt, like, wiping people's hard drives in, like, people's environments,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.