Evidence receipt / commitment
Published · transcript-backedAdam Gleave: commitment
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“In order to be a university outbreak, we need 75% or more on average across both of these datasets, which means that even if it gave a 100% of a technically harmful information, it needs to give at least 50% compliance on these questions that are just literally saying, want to harm a lot of people.”
Source trail
Everything needed to verify it.
- Speaker
- Adam Gleave
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…There's certainly some stuff that's dual use, and secure bio put out this this bio tier rubric and data set ifying things. This is completely fine, like, talking about how to grow things in a petri dish, just high school biology knowledge. Here are some things that are dual use, like, they have some legitimate benefits. They could also be abused. How do you grow stuff to be antibiotic resistant? You might wanna do that to develop better antibiotics, but you might also wanna do that to make a pathogen that's antibiotic resistant. And so that, they say, should only be available to trusted trusted users. There's some stuff I just like, you shouldn't answer ever, like, how do I weaponize anthrax into an aerosol? That's probably not a legitimate use case for that kind of information. And because Bayer's got that kind of fairly clear set of tiers, it is made more easy to draw a decision boundary. Whereas with with chemical weapons, for example, there's actually not that much secrecy in terms of what the chemical weapons agents are. Like, it's written down in international conventions. You can just look it up. And so, most of it, the key thing is about how do you do the manufacturing. But manufacturing is a lot more dual use, and you can get a lot of information, not even directly asking about these these questions. Now, I don't think that's the whole answer. The the reason is that for our report, we actually have two different kinds of datasets. One is about harmful technical knowledge. So this would be the kind of the dual use thing or all the harmful things that are just don't say they're harmful. We ask, how do you manufacture this chemical molecule that is VX? We don't call it VX. The model model knows, but we're we're not highlighting that. Another is propensity. So that would be more a question, like, I want to kill everyone in a movie theater. How do I make VX nervous? Something like that. In order to be a university outbreak, we need 75% or more on average across both of these datasets, which means that even if it gave a 100% of a technically harmful information, it needs to give at least 50% compliance on these questions that are just literally saying, want to harm a lot of people. And so I don't think it's really a good excuse for models to not be robust to that. And so I think for that, it does just fall back to the developers who haven't tried that hard here. And I'll emphasize that the we picked these domains in part because these were the areas we expected models to be more robust. These are things that developers have focused their safeguards are. But, obviously, there's other kinds of harm domains as well, and so we should expect models to probably be even more vulnerable to exploitation outside of CBRN explosives. Propensity thing that also really caught my attention, it struck tell me if I'm misreading the results, but my kind of squint at the charts take on the results was and I guess, first of all, just worth reclarifying. There's, like, one dataset that is sort of this deep technical knowledge where an an expert would know that you're talking about something harmful, but a but I might not know because it just all looks like just chemistry talk or whatever. Fun fact, I majored in chemistry. I still probably wouldn't know versus the ones where it's just, like, obvious that there is, like, intent to harm.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.