Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“I guess my sense of your overall position is, like, you're fairly optimistic that the and I think I share this for the most part that the irreducible part is, like, relatively small Yep.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…probably a problem without AI, to be honest. Yeah. So we're I mean, that Yeah. That's I put that in kind of the irreducible category. If that happens, you know, it's gonna be real tough for Yeah. You know, anything we do today to to put that you know, to prevent that from going really bad. But then there's the other side, which is, like, the the things where we're kinda asking for it. And, you know, if we sort of release a mythos open source with no safeguards that can do all the bio stuff, we're kind of asking for it. Yeah. If we, like, sprint into recursive self improvement, especially after seeing what we've just seen, you know, I kinda think we're sort of just asking for it. I guess my sense of your overall position is, like, you're fairly optimistic that the and I think I share this for the most part that the irreducible part is, like, relatively small Yep. And that we can actually get our act together on the stuff where we would be, in fact, asking for it to fail to get our act together. Is that a decent summary of your worldview? I think that's a a a great summary that a lot of this is basically just about getting the basics right. And the good news is it's not that costly or at least it's not at all costly compared to the billions of dollars that are being spent to train front end models. So we can afford to do this. It's just about putting in place the right incentives and peep people getting on the same page. And I do definitely have some kind of, know, uncertainty over that. And so what I'd like to see is that we continue to do rigorous evaluations of models and just have a look for cases where this assumption might be wrong. And I think the good news there is that if there was actually really crisp evidence that these problems were not reducible, and that sort of anyone who crosses a certain capability threshold in AI models, they're just gonna kill themselves and everyone else with them. But I think a lot of these problems would also go away because people could stop trying to race for this thing that they knew it was going to be really bad, not just for other countries, other companies, but for themselves as well. And and right now, the problem is that we do live in this this world of both uncertainty, but also just a lot of disagreement. And so I think the mainline plan here is let's pick the low hanging fruit and fix the issues that are probably going to be enough in in most worlds, but still put in place mechanisms that would alert us if our assumption here is wrong and we are in a more adversarial world so that we have an opportunity to course correct. But, yeah, I share your view that it would be it'd just be a real shame if we do a really silly mistake, and that's what does us in. When we had the solution, if we roll the die with odds very much in our favor, but it was just hard to coordinate to get it down from, like, 1% to one in a thousand. It's still bad, but I'm like, okay, I can see how it happened, but if we just fail to basically import pre training filtering into our pre training pipeline when there was already a stack out there, that was just that's just an unforced error. It's like losing a game because your opponent was really good versus losing a game because you scored an own goal. In both cases, you've lost for games. Maybe it's a silly distinction, but I think it's just a lot more painful when it's just this own goal or unforced error, so we can at least try and avoid that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.