High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Adam Gleave: prediction

30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

“security where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. So there, you do need more precision, but I expect developers are going to be tempted to give themselves quite a sort of wide berth around the areas that they don't have too much economic value behind and lean on trusted access programs for people the small set of people that do need access to those kinds of capabilities.”

— Adam Gleave

Source trail

Everything needed to verify it.

Speaker
Adam Gleave
Attribution
Verified speaker
Claim type
prediction
Recorded
30 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…plain text in the chain of thought. And I'd say that we've seen a real trend towards developers using probes. So these are small specialized models that are fit on top of activations of a main model. The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does. But, of course, the downside of this is that you're reducing the defense in-depth aspect. You're making your different defenses more correlated because they all ultimately rely on these activations. And if you can just fool the model's activations so this doesn't show up, then that no longer works. So this is the main safeguard stacks that we see terms of, like, hard refusals of a model not not directly answering requests. But we're also, of course, increasingly seeing developers rely on extra steps that could be having some kind of asynchronous monitoring of accounts so that, yeah, if you just keep on hitting these safeguards and you're trying to iterate on jailbreaks, then your account might get flagged and banned. Now I'd say that is a little bit early stage to really provide much assurance because you can just make new accounts. And we're seeing this happening at a sort of industrial scale. Bikini Anthropic, for example, has stopped Chinese based individuals and organizations creating Claude accounts. But everyone I know in China has a Claude So, you know, there's just reseller marketplaces. It's generally pretty hard to get to know your customer rate. Although, you could definitely imagine a a simple sort of dollar based amount where, like, you have to deposit $500 in order to be able to have your account be eligible for latest model. And then if you get banned before you spent the $500, suddenly this has become a much more expensive endeavor to try to abuse the system. So I think there are ways around it, but it's definitely not solved yet. And then I think the other interesting trend that we're seeing is, in some cases, developers is drawing a really big safety margin around the capabilities that they're worried about. This is actually quite annoying to users, including me. Like, I asked Fable five recently about how sake is fermented, and it said, no. This is a bio question. I'm gonna downgrade you to Opus. So if your refusal radius is so large to include, like, completely harmless bio questions, then you can see how you can make your model pretty robust against and we found no universal jailbreaks in in bio against harmful bio questions. But that's a lot harder for domains like cybersecurity where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. security where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. So there, you do need more precision, but I expect developers are going to be tempted to give themselves quite a sort of wide berth around the areas that they don't have too much economic value behind and lean on trusted access programs for people the small set of people that do need access to those kinds of capabilities. But then they're gonna have to work harder on the safeguards and getting the right decision threshold for these sort of large scale use cases that are ultimately generating their revenue. Is it is there do you know how much value each of these layers of defense in-depth provide? Like, how I guess you said, it's still pretty hard even when you just have model. So is it, like, percent is already caught in the base model and then you kind of incrementally, you know, get closer and closer to the goal with all these additional layers? Yeah. It's a hard…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence