High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Adam Gleave: evaluation

30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

“Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards.”

— Adam Gleave

Source trail

Everything needed to verify it.

Speaker
Adam Gleave
Attribution
Verified speaker
Claim type
evaluation
Recorded
30 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah. It seems like there's sort of a spectrum. Don't know where we're on it yet, and hopefully, we'll get the transparency that we've been promised. But I guess my kind of mental model of this right now is we're somewhere on a spectrum from extremely negligent to extremely scary. And on the extremely negligent end is like, had this thing. You gave it this task. You had no additional monitoring. You hope your sandbox is good enough. It turns out it's not. Next thing you know, you're totally pwned, but but it's not necessarily that scary because you're because you were in fact very negligent. Where on the other end, if you had, like, probes running and you had, like, additional, you know, reasoning models supervising the trace and this still happened, then we're, like, in extremely scary territory. Yep. I don't know if it's good or bad, but I I'm guessing we're more toward the negligent end. But I don't know. I mean, the vibe, you could you probably are I pride myself on being fairly plugged into the vibe. You're probably even, you know, significantly more plugged into the vibe. The vibe seems to be kind of, like, spooked. So it's it the vibe doesn't suggest that it was totally negligent. It does kind of my guess would be naively, you know, if I hadn't talked to people and seen how spooked people would be, I would think that this was, like, kind of, you know, bad practice, sloppy and negligent Yeah. But not so scary. But, like, people are acting like it's scarier than that would lead one to conclude. Yeah. Yeah. We just, as you say, don't have the full information. But I think from what we do know, it wasn't like OpenAI was just running some outdated piece of software, and the model downloaded a publicly available exploit and used it. Right? It did find a zero day in a widely used piece of software and used that to to break out of a sandbox. And you think you you could call it negligent, because maybe if you're playing with fire and you don't have a lot of containment, you're being negligent. But I'd be pretty surprised if OpenAI's internal infrastructure was less secure than your median American company or something like that. And suddenly, there's nothing in the report that suggests that would be the case, and there was also Hugging Face is a pretty capable company, and again, probably has much better information security posture than the median company around the world. So from that perspective, I think it it's it's right to be quite spooked that you're sort of run of the mill or even good, but not paranoid security practice is not enough to contain agents any longer. And maybe the extra AI specific control mechanisms will, because they're running in a harness, and you can inspect their actions. It's not like it's just someone who's got SSH access into your server. Maybe good, but not NSA level air gapped cybersecurity would be enough to stop AI agents from doing this, but for how much longer? Their cyber capabilities are getting better and better. And maybe to some extent, it it doesn't matter whether it's negligent or really scary capabilities, because this is evidence that there are gonna be actors who deploy with this level of safeguards. And I don't think that OpenAI is, by any means, the sort of most reckless actor here. And other developers are not that far behind. It's probably not more than three to six months, suddenly not more than a year behind OpenAI their capabilities. So this is something that needs to have some some kind of industry wide solution to ultimately, whether that be better cybersecurity across the board, easily deployed control mechanisms, or just more care and attention because people are aware of this kind of risk. That that's probably the main reason. I'm optimistic and not freaking out. It's like, okay. We're gonna see these warning shots. And so long as we respond appropriately, it's actually encouraging that we're seeing these issues, and it's not that an agent waits until it can definitely take over and then does a treacherous turn, which for some of us have, more more old school AI safety concerns. But I also see there have been a lot of warning shots, and at some point, I'm just wondering if we're gonna until something really bad happens. Yeah. It does seem like the vibe is, at least for now, taking this seriously. But yeah. Mean Yeah. The news cycle is short too. So who knows what will things will look like in even just a few weeks.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence