Evidence receipt / evaluation
Published · transcript-backedAdam Gleave: evaluation
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark.”
Source trail
Everything needed to verify it.
- Speaker
- Adam Gleave
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Now dealing with the nightmare scenarios, at least, you know, the early nightmare scenarios. How would you tell the story? I mean, everybody's heard the stories. I don't I don't mean in the basic sense of, like, we all know what happened. But how would you interpret the story? Like, what what's what's your, you know, unique way of contextualizing what we've just seen? Yeah. Maybe I'll start by just, I think, highlighting the things that aren't surprising about this story or aren't surprising if you've been following alignment in AI carefully. So reward hacking behaviors where models do really quite sophisticated cheating on tests. We've been seeing that for years, and people have been warning that's gonna happen for a decade. But in until now, these are being pretty toy, almost like cute misbehaviors over model edited its unit tests. So what? Now, have a model that is going out hacking third party computer systems to see if answers to a test. In some ways, from a model's perspective, this is no different to what it has been doing before. It's just it's got the capabilities to cause real harm. And so, I think that that should be a bit of a wake up call, but there's a lot of problems that have become almost mundane. We've seen them so many times before that we start saying, oh, yeah, this is an alignment failure, but does it really matter? They're gonna start mattering as the models get more capable. So, in a way, that's optimistic because we've got a pretty clear empirical set of warnings and things we can optimize and improve. And in some ways, it should be a bit pessimistic because we've had all these warnings and we've disregarded them, and it's not clear that we that's gonna necessarily, you know, change without some some external pressure. And then I think the part of the story that isn't talked about that much is that, at least as as far as we can tell from public information, Hugging Face noticed this hack first. So OpenAI was running an internal deployment for testing. It went out and hacked a third party. And it wasn't OpenAI having an automatic monitoring scheme and being like, oh my god, what has our agent done? We need to call up Hugging Face and apologize. Hugging Face noticed it. They investigated, and it turns out, oh, oops, it was an OpenAI model. So I'm sure OpenAI is going back over their transcripts now. You'd have to wonder, is there a chance it hacked anyone else? Or if it hadn't hacked Hugging Face, hacked someone who had a worse security posture. Would we ever have noticed? And it isn't the first time that something like this has has happened. Alibaba was training their Rome system, and their cloud security team called them up and said, hey, I we think that someone has compromised your training service because there's cryptocurrency mining going on. And it turned out that the agent had thought, oh, I I need to get some more compute. How do I do that? I'll I'll do some mining on the compute I have, and I'll try and rent some servers elsewhere. And in in some ways, I think that's even more of a near miss loss of control incident because it was actually trying to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark. to start gaining resources and potentially copy itself outside of infrastructure, whereas at least the OpenAI model have this pretty narrow objective of just getting some test results on a benchmark. So, yeah, I think takeaway from this would be less on the alignment side because in fairness to OpenAI, this was a model that had cyber safeguards removed, and I don't think we know of exact prompting regime, but it might well have been told to do something like this or at least incentivized to do it. So, it's not where we can't align these systems, but it is a massive control internal monitoring failure where the sandbox was insufficient. It doesn't seem like there was any additional layer of control mechanisms on what the AI system was doing. There wasn't asynchronous monitoring that alerted us us to that. So I think that really needs to change, both for prosaic reason that, you know, sooner or later, you're gonna hack someone who doesn't take this as gracefully as Hugging Face does, and your company is gonna be in a lot of trouble. And and also for the the risk that we're gonna see with more capable models that might be pursuing much more malign goals than just trying to cheat on a test.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.