High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Adam Gleave: evaluation

30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

“When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem.”

— Adam Gleave

Source trail

Everything needed to verify it.

Speaker
Adam Gleave
Attribution
Verified speaker
Claim type
evaluation
Recorded
30 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…If you zoom out and just kind of ask, okay. For the companies that are trying the hardest, are we which seems like it's kind of two right now, With all the techniques they have and all the techniques that they sort of, you know, that you've kind of mapped out there that they maybe haven't fully implemented yet, but, obviously, you know, it's especially the know your customer type stuff or the minimum deposits, like, that doesn't take a lot of technical wizardry, right, to to just implement that kind of basic blocking and tackling. Are we offense dominant, or are we defense dominant, over the next couple years? Yeah. I so I spaked a lot of my career actually arguing for this being offense dominant. I was very skeptical that we would solve adversarial robustness, and I've been walking about area for a decade. But I have to say the way wind winds are blowing, at least when it comes to LLM agents providing detailed multi turn assistance to harm for requests. Seems like it's defense dominant with the right technologies. And I I think that the reason for that is this defense in-depth approach. You don't just have to stop a model ever misclassifying something. You can have multiple different kinds of defenses from account level bans to externalized safeguards to model alignment, and it is increasingly hard to slip through all of those cracks persistently. But also, there there is this fundamental difference between the classic adversarial example setting you see in machine learning, like you add some white noise to an image and it flips the classification versus this kind of harmful assistance where you're not just flipping a classifier from, you know, one category to another. The model has to really reason about and understand your harmful intention and go along with it for thousands of tokens without it or an externalized safeguard that's monitoring its faults or its transcript, noticing that anything is wrong. And so that that's actually, fortunately, a much easier problem to stop, especially if you're willing to draw a bit of a a safety buffer around it and refuse some requests for that a dual use. So I think that's the optimistic take I have on it. To give up the pessimistic take, I would say that the the dual use part is actually quite challenging. And, I think we're seeing this with cyber security already. When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem. Right? So there's an offense defense balance in cyber, which relies on the defenders also getting access to these these capable models. And, unfortunately, a a lot of things in the world are just dual use. And I I think it would be a mistake for us to just point blank refuse on that. That that is bad for for the world. And you can get some way through trusted access programs and understanding before contact center is coming. But you are ultimately gonna end up in a situation where if you will allow dual use queries, and I think we need to allow a lot of them, you're gonna get some abuses. And so then that becomes a science or resilience question of if we're gonna have bad guys abusing models for cyber, how do we also really speed up the patch time? I've heard terrifying things that hospitals take in more than a year to update their operating systems. That's not gonna work in this environment. They need to be updating it within a a few days. And that, unfortunately, is just gonna be quite quite an expensive thing. to update their operating systems. That's not gonna work in this environment. They need to be updating it within a a few days. And that, unfortunately, is just gonna be quite quite an expensive thing. So maybe it is defense dominant for AI, but I don't know if a bad applications of AI, if those are offense or defense dominant. I'm optimistic that for cybersecurity, we can take this defensive acceleration approach and eventually just rewrite all of our software into memory safe languages and do formal verification. There's kind of all this stuff that AI could enable that's really good for the defender. But when it comes to something like bio, I don't think that we are going to be able to use AI to rewrite the human genome to be robust to viruses. At most, you might be able to speed up vaccine development, you've still got to manufacture a thing and get it in people's arms and run clinical trials. AI is gonna be able to have modest speed up some of these, but it's not gonna fundamentally change the physical reality. So that that's where I'm I'm more pessimistic that even though we can probably really hold back many of these things, it's ultimately gonna be more about buying us time to invest in societal safeguards rather than just being able to completely prevent misuse of models.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence