Evidence receipt / prediction
Published · transcript-backedNathan Labenz: prediction
6 Jun 2026 The Cognitive Revolution AI in the AM — Week 1 Highlights (June 2026)
“One of the interesting ideas that I heard there that I had not heard before was that the model that you would want to have internally for AI research might have quite a different constitution from the one that you deploy publicly for kind of general purpose AI assistant use cases. And they seem to think that, in fact, you probably would want to have something even more focused on safety and more sort of restricted in some ways, but maybe also less inclined to refuse certain tasks, but basically a different behavioral profile, which I do think is interesting because if you're going to make this sort of chain of thought monitoring plan work, I do think you're probably going to need some meaningful diversity of the AIs.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 6 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…r than they have moved and kind of pull away from the competition. I would say most people at that event thought that was very credible. There was not too much debate around like, will this level off? Now, obviously there's some selection effect there, but you could just go to, the whole event was under Chatham House rule, so I will respect that and not attribute specific statements to specific people or organizations. But you could go to the recursive website to look at speakers, whose identities were shared obviously with their permission. And, they've definitely got some notable people from the frontier companies. So these were not people that are fringe or, who you would say likely don't represent, kind of mainline views at the companies. It really seemed that the expectation is, yes, this is going to work. It's going to have a major accelerating effect. We don't necessarily know if it's going to have a sort of simple accelerating effect. Like in a human organization, if you went from 1000 to a million researchers, you probably wouldn't get 1000 X output. So there may be some sort of, you know, coordination challenges or just kind of duplication challenges that we see in human organizations. Maybe that happens in the same way. That's one possibility where you still get acceleration, but it's not a like, you know, blinding kind of takeoff acceleration. Or, you know, I would say also understood to be a credible, realistic possibility was that it is even a more profound phase change than that. And things like pre-training just become dramatically more efficient and models, suddenly have all these new qualitative abilities that they didn't used to have, such as, continual learning that really works or what have you. And so everything could change, in a very dramatic way, potentially very quickly once these milestones are hit. In the room, people said, and it did, there was quite a distribution. I was pretty much right at the median when we were asked, how many copies of you would it take to do the work that you are currently doing with the benefit of AI? The median answer was basically two. In other words, people felt like they're getting two times as much work done thanks to AI. But that was also framed in an interesting way where it was like, but note that as of today, at least, if you were not there, your productivity would drop to close to 0. Not too many people felt that they had any system that would continue to work in any sort of meaningful way if they were entirely removed from the picture. So there's a significant productivity boost, but there's still this sort of necessity of at least some human salt into the recipe to get the whole thing working. And then a big part of the discussion too was like, how can we set that up in such a way or create some sort of self-correcting structure or some sort of governance mechanism that can keep that on the rails, broadly speaking? By far, the number one strategy seems to be monitoring. h a way or create some sort of self-correcting structure or some sort of governance mechanism that can keep that on the rails, broadly speaking? By far, the number one strategy seems to be monitoring. It's very, we're very, very, as a civilization, whether we know it or not, listening to people at the Frontier Labs who are about to, in their own minds, and I believe they're probably right, set off this relatively uncontrolled experiment of AI recursive self-improvement. The big thing that they are betting on is AIs monitoring other AIs. It's very like monitoring the chain of thought, watching out for bed stuff, maybe training some different models. One of the interesting ideas that I heard there that I had not heard before was that the model that you would want to have internally for AI research might have quite a different constitution from the one that you deploy publicly for kind of general purpose AI assistant use cases. And they seem to think that, in fact, you probably would want to have something even more focused on safety and more sort of restricted in some ways, but maybe also less inclined to refuse certain tasks, but basically a different behavioral profile, which I do think is interesting because if you're going to make this sort of chain of thought monitoring plan work, I do think you're probably going to need some meaningful diversity of the AIs. Like we already hear from practitioners all the time that you want to have a model from a different model provider do the critiques because their failure modes are just a little bit different and you get better critiques, you find more issues that way. So they are thinking that way a bit internally, but they're very, they're very focused on this phenomenon, making it happen, figuring out some ways, hopefully to kind of keep it on the rails. I was honestly not that impressed with the quality of planning that we heard. y're very focused on this phenomenon, making it happen, figuring out some ways, hopefully to kind of keep it on the rails. I was honestly not that impressed with the quality of planning that we heard. It was very much like, We're going to try to figure it out as best we can. We're going to have AIs to help us. They will do a ton of monitoring. Like we're just going to pour compute on the monitoring side. And hopefully that will kind of work out for us. Also notably, there was a general kind of shared understanding that we might need to do some sort of coordinated slowdown at some point. Like the the sense that we might not be able to pull this off and that we, hopefully will recognize that and not just blindly, go off the cliff. There was, I would say, a remarkable amount of not just like cross lab camaraderie, because, I would say people are generally friendly to each other always, even if they're competing fiercely. But there was a sense that like, hey, we might need to really collaborate on slowing some things down if this phenomenon is starting to take off and our techniques aren't working as well as we might hope. So the open window in some way has shifted there, I think, where that is something people can talk about. There's also been this proposal recently of creating safe harbor for companies to cooperate on safety things where it might otherwise be considered an antitrust violation. And so I think that could be really good. I was pleased. I went in expecting basically to find that, or basically here, that yeah, we're like headed for this phenomenon. We have some ideas about how we're going to steer it in the right direction. And I didn't think I would hear that many great ideas. In fact, what I heard was even less compelling than what I expected. So I was sort of negatively updated in terms of the quality of plans people have. but positively updated in terms of their recognition of how inadequate the plans are and sort of their willingness to entertain that they might need to sort of break the frame of the race that they're currently running against one another in order to just again, not blindly race off the cliff. So I thought that was good. Then I tried something I just watched those same lab leaders agree on stage that the AI should do and went looking for why it wouldn't. But it was striking at the recursive event how just how few AIs people seem to think there really are going to be. And the disconnect, there was one panel discussion, careful to speak about this in the Chatham House rules abiding way, where people from multiple frontier model developers were speaking about their different approaches. And obviously, Anthropic is associated with the constitutional approach and open AI people are much more associated with the, you know, this thing should just follow the rules that we give it approach. And that's all public and certainly was not like a secret revealed at the event. But it was striking that like on one particular example that came up, which was AI helping people with a cigarette business, everybody agreed that the AI should do that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.