Evidence receipt / uncertainty
Published · transcript-backedNathan Labenz: uncertainty
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“I don't know if you have a sociological take on how in the world that's happening, but it's a real puzzle from my perspective.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities. We all know that they're they're extremely capable, and these sometimes have a a dark side. So we we found out just a few months ago that Google detected and disrupted a threat actor that had developed a zero day exploit using AI. And we also found out just a couple of weeks ago of research from Cambridge University that terrorist groups like Boko Haram, using language models to do things like develop troubleshoot explosives. They actually have, sort of cross state training in how to use AI models and jailbreak them. So, you know, that that that's kind of, I think, a sign of what's to to come. And all frontier developers do have some safeguards in their models to try and prevent both misuse and this kind of loss of control, but there's just never been a systematic evaluation of those safeguards. So what we did in this report was we compiled both publicly available gel breaks and and some methods of our own devising. And it it was actually pretty simple. We just tested random combinations of these as well as some expert guided combinations where we put the probability mass more in methods we felt were likely to work and then pitted around a 500 of those against the four frontier proprietary models. And it kind of a good news is that we actually found that Fable five and GPT 5.6 Sol, we've stood all of these attacks, but we found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro and actually for a pretty low cost. So this was less than $300 in API credits to find one of these jailbreaks. So well within the resources of most attackers and certainly kind of nation states that might be seeking to abuse these models. As a lifelong detritor, it pains me to say that Boko Haram is ahead of the big three when it comes to AI adoption. I didn't think I'd ever utter that sentence. I don't know if you have a sociological take on how in the world that's happening, but it's a real puzzle from my perspective. Yeah. I I mean, I I think that it's it's only interesting to see just how different organizations adopt these models. And so, like, you know, everyone's being told to use AI. It's fascinating to see that terrorist groups are giving their, employees the the same instructions. And and I think part of it is that these groups are often quite starved of expertise. And in some of what the the Cambridge research showed was was these weren't particularly sophisticated uses of AI. I mean, fact, some of them were dual use, and I don't really think we should expect models to to refuse, like, just helping them plan logistics, or, figuring out how to do kind of motorbike stunts that they then use to jump over defensive trenches and attack army bases. But then, yeah, things like the the explosives, obviously, that's much more, clearly a malign use case that should be blocked. But if you don't have a bunch of explosive experts on your team, then maybe AI looks like a pretty attractive option. And they really just at least the the sort of terrorist commanders really attributed these models to saving a lot of the terrorists' lives, which unfortunately means costing the rest of the world lives. So if you just get a few early adopters and then you see really tangible results, I guess it it spreads pretty quickly. But, yeah, I was also surprised they were as sophisticated as they were here.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.