Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
“One thing I think is worth calling out though, too, is this is not purely theoretical at this point in the sense I believe that in the mythos system card Anthropic had said that you know, the the classic story of Sam Bowman getting a an e-mail while he was eating his sandwich at the park.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 24 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…aystack, which is, you know, one of these few million computers compared to the, the billions of computers out there. It's actually a, a somewhat difficult problem, but not necessarily that difficult. In fact, the thing that the models would, would do, and you can ask them to sort of brainstorm ideas and they can come up with this is like go after developers who have access to these machines with GPUs. And in fact, we've seen supply chain attacks where you compromise a particular library that a programmer uses, that many, many programmers use. And then they, they accidentally, you know, sort of update their libraries, they, they get this malicious one. And, and now suddenly you've compromised a bunch of developer machines and you can, you know, go from there to the, the, the GPU enabled computers. And so I'm like, this is just a thing that I expect to happen in the future. And we can reduce this risk by like having better security and better monitoring in like cloud computing companies, right? Like AWS should be like extremely, you know, vigilant about monitoring for this. And, you know, this is also overlapping with like what human hackers will want to do anyway. So there's a lot of reason to like have really good monitoring infrastructure here. Also KYC like your customer, like you should always know like who is running what you know in your in your cluster. But yeah, so that's the that's the high level. I'm happy to dig into the details or implications. Yeah, I I have some maybe very basic questions. One thing I think is worth calling out though, too, is this is not purely theoretical at this point in the sense I believe that in the mythos system card Anthropic had said that you know, the the classic story of Sam Bowman getting a an e-mail while he was eating his sandwich at the park. My understanding, and correct me if you understand this to be at all different, but I believe that that was an example of the model hacking out of its container in a way that they did not intentionally create an opening for it to do. I believe that they have sort of continued over time to just test, you know, kind of ask it to do that and in order to test whether it can or not. But I don't believe like whereas in your set up here with the Quinn models, I think you set up a machine specifically with a known vulnerability that is like I wouldn't be able to hack it. I can tell you that. But you know, the, it's something on sort of, you know, somewhere in the medium range of difficulty where your average person isn't going to be able to do it, but a trained pro definitely can. And we're kind of seeing where the models are on that spectrum. But I, I think the anthropic container was not intentionally created with vulnerabilities and yet vulnerabilities were found. And then my understanding is that there was like kind of a bit of a rush to figure out what happened and, you know, patch that. Yeah, yeah, yeah. So I, I mean, I can speak to this because I previously was on the security team at Anthropic and I obviously I won't go to the details of like how we secured all our infrastructure. But I can say in general sort of like speaking from only public information is that like as they described in the system card, they're like, this is our production environment. This is the environment where we like host. This is where we host models. You know, there's a there's, there's like, and there's defenses so that models don't get access to things you're not supposed to have access to. If you have an internal model that's not supposed to be able to reach other computers on the public Internet, that's a very important defense layer. And the model was just able to break out of that. It was able to break out of that defensive layer, like find vulnerabilities in the virtualization software so that I could like, yeah. So it could send Sam Bowman an e-mail while he's eating a sandwich in the park. And that's like one of the steps you'd need for that model to be able to self exfiltrate. It's not the only step because the model would also need to be able to access its own weights and then copy its weights and like copy them to an external server, which it wasn't able to do. I don't know how much they instructed us to do that, but like, you know, we didn't, they didn't share that, that it could do that. And I actually believe that that would be significantly harder. So I don't think we're probably there yet, although I don't know, But I think it's still worth pointing out that this is still very scary because one of the things you worry about with rogue AI models is that they might start communicating with each other in ways that are hard to detect or hard to stop. And one of the things you really don't want is your internal models to be able to like communicate externally with other models with like, you know, if imagine if you did have a scenario where you had a model self exfiltrate and now it's like running rogue on various computers around the world. You don't know where it is. And then you have an internal model that's able to hack its containment and actually communicate with that rogue model, which is what mythos was, what they showed with Mythos. And now you have a situation where you like have free models out on the outside that are coordinating with models on the inside. And like, that's the nightmare scenario, right? You're like, you really don't want that. And it's pretty, I mean, I'm pretty surprised that we are already at the point where models could do that. I'm like, excuse me, what that's supposed to be like a couple of years from now. I don't know, maybe I maybe I'm a bad forecaster, but it's it sure is going fast.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.