← All source episodes The Cognitive Revolution / episode intelligence
Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
30 Jul 2026 16 published claims 2 attributable people
Speakers in the public record
Claim mix
evaluation 7prediction 4belief 3uncertainty 1commitment 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
16 published records
“Yeah. It's it's it's it's getting real, and I think, you know, that that's ultimately a big part of a motivation behind this security leaderboard is that, a lot of attention is paid to model capabilities.”
- Publisher
- The Cognitive Revolution
“I don't know if you have a sociological take on how in the world that's happening, but it's a real puzzle from my perspective.”
- Publisher
- The Cognitive Revolution
“Other countries can also be relevant if we're talking about more than a year of delay. So I think mostly, I'm operating under assumption that things are gonna continue to march on, that we can pick some of the low hanging fruit here on coordination and at least avoid the sort of most extreme race to the bottom on safety.”
- Publisher
- The Cognitive Revolution
“I guess my sense of your overall position is, like, you're fairly optimistic that the and I think I share this for the most part that the irreducible part is, like, relatively small Yep.”
- Publisher
- The Cognitive Revolution
“Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning.”
- Publisher
- The Cognitive Revolution
“I expect open weight developers to be the first to have to adopt this because they have fewer options.”
- Publisher
- The Cognitive Revolution
“I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark.”
- Publisher
- The Cognitive Revolution
“security where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. So there, you do need more precision, but I expect developers are going to be tempted to give themselves quite a sort of wide berth around the areas that they don't have too much economic value behind and lean on trusted access programs for people the small set of people that do need access to those kinds of capabilities.”
- Publisher
- The Cognitive Revolution
“In order to be a university outbreak, we need 75% or more on average across both of these datasets, which means that even if it gave a 100% of a technically harmful information, it needs to give at least 50% compliance on these questions that are just literally saying, want to harm a lot of people.”
- Publisher
- The Cognitive Revolution
“There's kind of all this stuff that AI could enable that's really good for the defender. But when it comes to something like bio, I don't think that we are going to be able to use AI to rewrite the human genome to be robust to viruses.”
- Publisher
- The Cognitive Revolution
“When OpenAI's testing agent went rogue and hugged Haggingface, they they had to use an open weight model to analyze it because the the closed weight models refused to help them on the defense side, and and that's a real problem.”
- Publisher
- The Cognitive Revolution
“The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does.”
- Publisher
- The Cognitive Revolution
“You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training.”
- Publisher
- The Cognitive Revolution
“And I always struggle a little bit operationalizing question, but I'm probably somewhere near 10% existential risk in the next few decades from AI. And and I think that we could probably get that down to something like 1% without any major research breakthroughs, just iterating and refining what we already have and taking a sort of careful engineering approach to systems and having good safety cultures at companies.”
- Publisher
- The Cognitive Revolution
“Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards.”
- Publisher
- The Cognitive Revolution
“Because what we found is that even though we can usually jailbreak the model, it's really hard to get it to shut up about the evil thing that it's about to do when it's reasoning in the chain of thought. And this is where the externalized safeguards can come in because if you have a specialized model that's looking at the input, the chain of thought, the internal reasoning of a model, the output, and trying to block compositions that kind of go in the the wrong direction, it that can be quite hard to bypass, especially when the model is thinking in-depth about it because you can normally get the model to obfuscate its outputs, and you can obfuscate the input.”
- Publisher
- The Cognitive Revolution