High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Adam Gleave: prediction

30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

“And I always struggle a little bit operationalizing question, but I'm probably somewhere near 10% existential risk in the next few decades from AI. And and I think that we could probably get that down to something like 1% without any major research breakthroughs, just iterating and refining what we already have and taking a sort of careful engineering approach to systems and having good safety cultures at companies.”

— Adam Gleave

Source trail

Everything needed to verify it.

Speaker
Adam Gleave
Attribution
Verified speaker
Claim type
prediction
Recorded
30 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…I mean, would you I guess in terms of, like, why this happens, just did a episode not too long ago with David Odd, David Dalrymple, who basically says, you know, he used to have a Pdoom. He would quote 70 plus percent. Now he's down to under 5%. Wow. That's a big change. Why? Well, basically, constitutional alignment seems to be working, and we can sort of, you know, get these AIs to become Bodhisattvas and, you know, work for the benefit of all living beings. And we just need and his kind of, you know, caveat is, like, we just need to not turn the RL up so high. We keep doing that because it obviously does bring performance. But, you know, every time we tilt too far toward RL, we start to see, like, o three type problems and, you know, little, you know, like, on end and bench type things. You know, you see Claude kind of doing things that they probably don't want Claude to be doing. So I guess one maybe overly naive story, but perhaps captures, you know, a decent chunk of what's going on is, like, don't overdo the RL. You know? Do you think that that could we draw a lesson as Yeah. Simple as that? And if we can, can we operationalize it into some rules? Yeah. I I think this is pointing out something really important, and maybe my my p team is not that different to to David Eds. And I always struggle a little bit operationalizing question, but I'm probably somewhere near 10% existential risk in the next few decades from AI. And and I think that we could probably get that down to something like 1% without any major research breakthroughs, just iterating and refining what we already have and taking a sort of careful engineering approach to systems and having good safety cultures at companies. And not like stopping AI or anything, but if we don't know how to make the next system align, then maybe we do do experiments and do more detailed evaluations before we deploy the next system. You know, just gating deployments on on actually doing that that rigorous work. It's still gonna look from the outside like things are moving incredibly fast, probably faster than they are now, but we're not just literally moving as quickly as possible, I think. You know, if you're really trying to move as quickly as possible, then there's gonna be a big pressure to cut corners on these things. I don't know if it's gonna be as simple a recipe as, you know, just don't do more of an x percent RL. And ultimately, some types of RL training are are really important in the models having a a line personas. I think it depends. Is this coming from human feedback? Do you have a human in the loop? Is this coming from constitutional AI? How much do you trust the constitution? Do you have some evaluator? Is this RL verified reward where you're just optimizing the system to solve certain tasks? Are you accidentally training on the chain of thought, which some developers do now and again and are meant to not do? So there's all these kinds of implementation details, but you could definitely imagine something that is operationalized as, okay, we know this training recipe seems to basically work and the models are pretty aligned. We know that if we push in this direction, the models become more water hacking. We think this is a dangerous part. We're gonna set a safety margin of don't go beyond this point. And then we're periodically gonna reassess if we've crept into the safety buffer or we're set more conservative thresholds for the next model. But if we've got various different kind of advances to our training stack, actually, we're quite, you know, way away from the safety margin because we've got these better techniques where maybe we can crank up the, you know, RLVR or something like that. I think that would be a great regime to to live in. It doesn't seem that hard technically. It's maybe quite hard economically and politically. And right now, there's a really huge amount of distress between the AI companies. Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning. een the AI companies. Like, I I was at a a workshop we ran recently on on chain of thought monitorability, which I think is maybe an easier one to operationalize than this this reinforcement learning. So we've got this kind of gift right now that that models, they just talk in their chain of thought about how they know that in an evaluation or are deceptive. And this is really useful from a capability standpoint because you can see why the model's making a mistake, and you can maybe trick tweak your post training pipeline. It's also really useful from a safety standpoint because you can see how the model might be trying to deceive you. But there's various ways we might lose this. We might train against a chain of thought accidentally. We might train models that have reasoning, neuralese, these continuous activations rather than chain of thought. So there was this proposal of could we just agree not to train neuralese models without letting other developers know? So it's not that you can't train it. Just give everyone a heads up. And this was even that kind of for people, were like, oh, it's gonna be really hard. That's information intellectual property leakage. We're telling people we're doing this. We're not telling people how. We're just saying that we're gonna do it. It's one bit of information. But even that, a lot of developers are really reticent about. I think it's good to try and work on things like this for sure, but I suspect there is gonna need to be some kind of third party that comes in and addresses this. I don't think it's gonna be easy for the developers to do that themselves.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence