High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Adam Gleave: evaluation

30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

“Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities.”

— Adam Gleave

Source trail

Everything needed to verify it.

Speaker
Adam Gleave
Attribution
Verified speaker
Claim type
evaluation
Recorded
30 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…I'm excited. This is obviously an increasingly critical moment in AI history. I think it's that's the nature of exponentials. It kinda keeps happening that way, and it's probably gonna continue for a little while to come. So, also, the occasion for this conversation is that FAR has just put out an AI security leaderboard. And so we're gonna start by kinda digging in on that and understanding the details of of that work, why you're doing it, what it why it matters, why it matters is pretty obvious, but I'm really interested to get into some of the nitty gritty. And then also to zoom out and kinda take stock of where we are as we've now got legitimate breaking out, loss of control, lab leak type scenarios coming into the real timeline that we're in. What a what a time to be alive. Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities. We all know that they're they're extremely capable, and these sometimes have a a dark side. So we we found out just a few months ago that Google detected and disrupted a threat actor that had developed a zero day exploit using AI. And we also found out just a couple of weeks ago of research from Cambridge University that terrorist groups like Boko Haram, using language models to do things like develop troubleshoot explosives. They actually have, sort of cross state training in how to use AI models and jailbreak them. So, you know, that that that's kind of, I think, a sign of what's to to come. And all frontier developers do have some safeguards in their models to try and prevent both misuse and this kind of loss of control, but there's just never been a systematic evaluation of those safeguards. So what we did in this report was we compiled both publicly available gel breaks and and some methods of our own devising. And it it was actually pretty simple. We just tested random combinations of these as well as some expert guided combinations where we put the probability mass more in methods we felt were likely to work and then pitted around a 500 of those against the four frontier proprietary models. And it kind of a good news is that we actually found that Fable five and GPT 5.6 Sol, we've stood all of these attacks, but we found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro and actually for a pretty low cost. So this was less than $300 in API credits to find one of these jailbreaks. So well within the resources of most attackers and certainly kind of nation states that might be seeking to abuse these models. As a lifelong detritor, it pains me to say that Boko Haram is ahead of the big three when it comes to AI adoption. I didn't think I'd ever utter that sentence. I don't know if you have a sociological take on how in the world that's happening, but it's a real puzzle from my perspective.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence