Evidence receipt / evaluation
Published · transcript-backedAdam Gleave: evaluation
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does.”
Source trail
Everything needed to verify it.
- Speaker
- Adam Gleave
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. So all of the developers that we've, you know, we we've looked at do have some kind of defense in-depth. But so as you allude to, how deep it is and what those individual components is quite different between developers of where we're starting to see some convergence. So we just map out the different pieces. The most basic is the model itself that you're interacting with as generating these responses. All of them have undergone some kind of post training, both for instruction following, but also for some kind of refusal training. And it really varies in how sophisticated that that is. In some cases, it's a sort of pretty basic pipeline with static examples of what to say yes or no to and some human feedback and a lot of verified reward environments to make it really good at programming, but does very little adversarial robustness. In the other extreme, some developers have gone all the way to large amounts of synthetic data generation. And different developers have different approaches, and Anthropic has got this constitutional AI approach that has another model score things according to a rubric. Open AI had this adversarial training approach using self play where they trained another model to basically be a red teamer and find vulnerabilities and then trained the model to be robust about and then trained for red teamer to be better at it. And I think both of these are pretty good approaches. Ideally, people would combine all of these. And so it it definitely even when we've had access to no safeguard models or no no external safeguards, just the main model itself, it's definitely gotten hard to jailbreak some of the best models. So that that this is an important part of the the pipeline. But, actually, in some ways, important part is making sure that the model kind of verbalizes what it's doing. Because what we found is that even though we can usually jailbreak the model, it's really hard to get it to shut up about the evil thing that it's about to do when it's reasoning in the chain of thought. And this is where the externalized safeguards can come in because if you have a specialized model that's looking at the input, the chain of thought, the internal reasoning of a model, the output, and trying to block compositions that kind of go in the the wrong direction, it that can be quite hard to bypass, especially when the model is thinking in-depth about it because you can normally get the model to obfuscate its outputs, and you can obfuscate the input. You just simple things like shifting rot 13, shifting every letter halfway through the alphabet was enough to get through early models. You gained more sophisticated techniques now, but models are perfectly able to read all sorts of obfuscated inputs. But, usually, they'll just reason in plain text in the chain of thought. And I'd say that we've seen a real trend towards developers using probes. So these are small specialized models that are fit on top of activations of a main model. plain text in the chain of thought. And I'd say that we've seen a real trend towards developers using probes. So these are small specialized models that are fit on top of activations of a main model. The main thing driving this is compute efficiency because you can train a lot of probes and run them at deployment time with minimal overhead because you've already computed all of activations of a main model. And it does have a benefit that it is able to use all of our internal representations of a main model, which is usually quite powerful because one of the problems with having smaller specialized language models is the the filters is that you might be able to do some obfuscation scheme, but they don't understand that the main model does. But, of course, the downside of this is that you're reducing the defense in-depth aspect. You're making your different defenses more correlated because they all ultimately rely on these activations. And if you can just fool the model's activations so this doesn't show up, then that no longer works. So this is the main safeguard stacks that we see terms of, like, hard refusals of a model not not directly answering requests. But we're also, of course, increasingly seeing developers rely on extra steps that could be having some kind of asynchronous monitoring of accounts so that, yeah, if you just keep on hitting these safeguards and you're trying to iterate on jailbreaks, then your account might get flagged and banned. Now I'd say that is a little bit early stage to really provide much assurance because you can just make new accounts. And we're seeing this happening at a sort of industrial scale. Bikini Anthropic, for example, has stopped Chinese based individuals and organizations creating Claude accounts. But everyone I know in China has a Claude So, you know, there's just reseller marketplaces. It's generally pretty hard to get to know your customer rate. Although, you could definitely imagine a a simple sort of dollar based amount where, like, you have to deposit $500 in order to be able to have your account be eligible for latest model. And then if you get banned before you spent the $500, suddenly this has become a much more expensive endeavor to try to abuse the system. So I think there are ways around it, but it's definitely not solved yet. And then I think the other interesting trend that we're seeing is, in some cases, developers is drawing a really big safety margin around the capabilities that they're worried about. This is actually quite annoying to users, including me. Like, I asked Fable five recently about how sake is fermented, and it said, no. This is a bio question. I'm gonna downgrade you to Opus. So if your refusal radius is so large to include, like, completely harmless bio questions, then you can see how you can make your model pretty robust against and we found no universal jailbreaks in in bio against harmful bio questions. But that's a lot harder for domains like cybersecurity where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. security where a lot of legitimate use cases that are ultimately making these companies money through coding agents look very similar to the kinds of offensive or at least dual use cyber capabilities. So there, you do need more precision, but I expect developers are going to be tempted to give themselves quite a sort of wide berth around the areas that they don't have too much economic value behind and lean on trusted access programs for people the small set of people that do need access to those kinds of capabilities. But then they're gonna have to work harder on the safeguards and getting the right decision threshold for these sort of large scale use cases that are ultimately generating their revenue.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.