Evidence receipt / belief
Published · transcript-backedAdam Gleave: belief
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning.”
Source trail
Everything needed to verify it.
- Speaker
- Adam Gleave
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. I I think this is pointing out something really important, and maybe my my p team is not that different to to David Eds. And I always struggle a little bit operationalizing question, but I'm probably somewhere near 10% existential risk in the next few decades from AI. And and I think that we could probably get that down to something like 1% without any major research breakthroughs, just iterating and refining what we already have and taking a sort of careful engineering approach to systems and having good safety cultures at companies. And not like stopping AI or anything, but if we don't know how to make the next system align, then maybe we do do experiments and do more detailed evaluations before we deploy the next system. You know, just gating deployments on on actually doing that that rigorous work. It's still gonna look from the outside like things are moving incredibly fast, probably faster than they are now, but we're not just literally moving as quickly as possible, I think. You know, if you're really trying to move as quickly as possible, then there's gonna be a big pressure to cut corners on these things. I don't know if it's gonna be as simple a recipe as, you know, just don't do more of an x percent RL. And ultimately, some types of RL training are are really important in the models having a a line personas. I think it depends. Is this coming from human feedback? Do you have a human in the loop? Is this coming from constitutional AI? How much do you trust the constitution? Do you have some evaluator? Is this RL verified reward where you're just optimizing the system to solve certain tasks? Are you accidentally training on the chain of thought, which some developers do now and again and are meant to not do? So there's all these kinds of implementation details, but you could definitely imagine something that is operationalized as, okay, we know this training recipe seems to basically work and the models are pretty aligned. We know that if we push in this direction, the models become more water hacking. We think this is a dangerous part. We're gonna set a safety margin of don't go beyond this point. And then we're periodically gonna reassess if we've crept into the safety buffer or we're set more conservative thresholds for the next model. But if we've got various different kind of advances to our training stack, actually, we're quite, you know, way away from the safety margin because we've got these better techniques where maybe we can crank up the, you know, RLVR or something like that. I think that would be a great regime to to live in. It doesn't seem that hard technically. It's maybe quite hard economically and politically. And right now, there's a really huge amount of distress between the AI companies. Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning. een the AI companies. Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning. So we've got this kind of gift right now that that models, they just talk in their chain of thought about how they know that in an evaluation or are deceptive. And this is really useful from a capability standpoint because you can see why the model's making a mistake, and you can maybe trick tweak your post training pipeline. It's also really useful from a safety standpoint because you can see how the model might be trying to deceive you. But there's various ways we might lose this. We might train against a chain of thought accidentally. We might train models that have reasoning, neuralese, these continuous activations rather than chain of thought. So there was this proposal of could we just agree not to train neuralese models without letting other developers know? So it's not that you can't train it. Just give everyone a heads up. And this was even that kind of for people, were like, oh, it's gonna be really hard. That's information intellectual property leakage. We're telling people we're doing this. We're not telling people how. We're just saying that we're gonna do it. It's one bit of information. But even that, a lot of developers are really reticent about. I think it's good to try and work on things like this for sure, but I suspect there is gonna need to be some kind of third party that comes in and addresses this. I don't think it's gonna be easy for the developers to do that themselves. I should say just to represent Davidad correctly, he was specifically focused on RLVR and not, you know, all Sure. Manner of RL. That does seem to be where the real, you know, persistent increasingly aggressive behavior tends to come from. So that's a sobering note, though, on the kind of prospects for coordination. I wanted to ask you for your thoughts on prospects for US China coordination as well. But, you know, it's it's tough if we can't get our own house in order. Yeah. What do you think I mean, we are in this kind of interesting moment where, like, the Overton window has been blown wide open. Right? I mean, I this has been a the a quite a shift from not that long ago where all of a sudden there is people all over the place expressing openness to some sort of coordinated slowdown or whatever. How do you think that is best realized? And and how critical is it that we get China on board? And and what are your how optimistic are you that we might be able to actually pull all that off and kind of jog into the singularity instead of, you know, full on sprint into it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.