Evidence receipt / commitment
Published · transcript-backedNathan Labenz: commitment
5 Aug 2026 The Cognitive Revolution Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“Because that's gonna bring about all sorts of deceptive behavior, and it's such such an adversarial environment. How about we agree that for the next six months, we won't do that?”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 5 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…ason we keep falling back on, okay, how much compute? You know, how many chips, you know, can we use for this purpose? Because, like, training is a distinct thing that, like, is very easy to identify. And for other purposes, we kind of can't limit you that much necessarily. You could also potentially limit I'm spitballing how much compute you would get from unreleased models. Similarly, to try and contain that, because, like, if you have to go for the release process, then, like, the way that AI goes truly ballistic would be, like, you would, you know, you would use n to train n plus one, to train n plus two, train n plus three, to train n plus four. And if you were doing this, like, cycle became a month and then became a week. Then that's when suddenly, like, holy shit. But, like, if you had a rule that, like, you have to release these models or you've only used so much compute for inference on those models, then, you know, you have to go through the process of submitting this thing and releasing it and then exposing it to the public and giving us an idea of what it is and, like, allowing us to react to that information. And that then slows you down from going truly, like, realistic, but this has, you know, obviously, very limited effect on the amount of diffusion and the amount of other progress that you can make. And so maybe that's a good trade, but, again, that's the like, I thought about this for thirty seconds right now. I haven't been focusing much on the exact limitations, but, like, I do think that if you paste the frontier, you would you would be doing it by placing restrictions on methods of training. And that what you what you could do to develop and use internal unreleased frontier style, like, models that, like, actively recurs these methods. You would prevent, like, one day utility. Right? It's like all Yeah. I've had some similar ideas about just the relationship between unreleased models and released models. I do think there's something there that would be really helpful for preventing runaway internal and largely invisible processes. I take your point on not wanting the government to come in and say what training techniques can and can't be used, But I do wonder if there are ways and, they could be short term, but what if the two companies got together and said, it probably wouldn't be a good idea for us to train models with an RLVR on, like, how much money they make in the economy. Because that's gonna bring about all sorts of deceptive behavior, and it's such such an adversarial environment. How about we agree that for the next six months, we won't do that? Do you have any hope for those kinds of just very specific, we've identified a bad thing? It would obviously be really economically valuable, but we both know in our hearts it's probably not a good idea, and so we won't do it for a while at least. I got some hope of informally people talk and people realize these things are bad, and they're like, we're not gonna but also, I hope they're doing that now without needing to an agreement. OpenAI clearly is doing things that are of the Buddhist of the Buddha nature of train on making the most money on the internet. Like, they're not doing that literal thing. I don't think they're that foolish. And I don't think it's that easy to do it that way. But clearly, they are making that fundamental mistake somewhere in their training process with a probability of 9.9 something. And so they need to fix that. But, yeah, I I think that they are definitely, like, trying not to do the maximally stupid things. And they're aware of the maximally stupid. But, like, again, like, it's very easy to have things that, like, you are trying not to do creep into your training process and end up happening anyway? You could say, we're not going to have training processes that have reward misaligned behaviors. How do you do that? The way that you do that is you get your act together, and you are very, very careful and prudent about all of your RL environments and all of your different training problems. And you keep an eagle's eye out for this stuff on many levels, and you make sure that it works. There's no specific thing you can say. I mean, in theory, could have a thing of like, oh, if I just we agree that if we're gonna have these various cross tests, and we're gonna, like, check each other's environments, and, like, we will throw out any environment that, like, anybody identifies problems with. And if we discover an AI has been trained on these environments, we're gonna reach we go and want rewind to a previous checkpoint before that happened and start again, so we'd better be really careful because otherwise, we're gonna waste a lot of resources. Can do these things, but, like yeah. Yeah. Look. It's it's really tough. And, obviously, one one potential advantage is if we really are in a two horse race at this point, where the two horses are pretty compatible with each other and talk to each other and they kind of understand each other pretty well, then it becomes pretty reasonable for them to make a decision that, like, not entirely opaque to each other of how much they're going to just, like, be more prudent even if this means that the action goes somewhat slower. Because, like, that also just again, my my my fundamental belief is that within a year and almost certainly probably even six months, if you invest more in fixing these problems and addressing these problems and, like, being more prudent about these problems, you will end up with a model that has better utility in the marketplace than the person who didn't do that, even if you are trading off those resources against your your capability developer. I think that people are just making a mistake.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.