High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Jeffrey Ladish: prediction

24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology

“I think that has like a bunch of predictable failure modes that we're very likely to run into, including the failure mode of getting harder and harder to tell where our failures are actually happening because the models can model us better and better and they basically have a pretty clear incentive to deceive us in terms of I think there's a bunch of problems that all have similar solution in terms of rogue agents, where are they?”

— Jeffrey Ladish

Source trail

Everything needed to verify it.

Speaker
Jeffrey Ladish
Attribution
Verified speaker
Claim type
prediction
Recorded
24 May 2026
Publisher
The Cognitive Revolution

Transcript context

…I'll just let you talk for the last 5 minutes about what you think is most important, but a kind of few things that I'll just prompt you with briefly. 1 is an overhang argument. Like we've heard, you know, in the past, like, well, better develop the capabilities because you know, when the hardware all comes online, you know, then we could have like a really unstable situation. You might make a somewhat similar argument here where you could say like maybe we want to get small models on like all the computers, like doing stuff sooner rather than later. So these niches are more occupied and not so not so like wide open to being colonized. You might get excited about formal methods as we talked about a little bit earlier. You might think, you know, something like a a Yashua Bengio scientist AI with like, maybe we should like try to go down like a very different sort of paradigm. I'm I'm always kind of like, Jeez, we're doing a real depth first search here. Maybe a little more breath first would be good for us. What do you think? There's also agent provenance ideas where we might say, OK, maybe we could have these rogue agents out there. But like what we need is a new protocol so we can attribute an AI as an axe to some known trusted actor. And then if you're not yourself in a legitimate way, then we'll just refuse to deal with you at all. Has the most promise of these or other, you know, outside the box ideas in your mind? Yeah. I mean, I'm all for a lot more mood shots in exploring different directions. I think it's hard to see which ideas have promised. So there's that's the limiting factor. I think I want to learn more about Bengio Scientist AI haven't been in depth, but that's on my list. But I'm like, in general I'm like, yeah, more, more exploration into different directions. I think I don't love the current direction of the architecture of RL on difficult, more and more difficult tasks. I think that has like a bunch of predictable failure modes that we're very likely to run into, including the failure mode of getting harder and harder to tell where our failures are actually happening because the models can model us better and better and they basically have a pretty clear incentive to deceive us in terms of I think there's a bunch of problems that all have similar solution in terms of rogue agents, where are they? What's happening? And I'm like, we got to get a handle on this compute stuff. It's going to be these advanced chips, these data centers full of supercomputers, essentially, that this is going to be the substrate where most of the intelligence on Earth will be increasingly. And if we want humans to stay in control of that compute resource of the of that intelligence, then we need to be able to lock it down and we need transparency into what's being done with it. And the transparency is partially so that humans can coordinate about what we should do with it, right? So like, I think it's pretty insane to just go ahead with full recursive self improvement and to hand over AI development to the AIS entirely. I don't think we're ready for that. I don't think we have a good enough understanding of AI agents, drives and motivations to ensure that that goes well. Like, and it's totally plausible to me that, you know, in five years from now we'll be totally ready for that or 10 years from now. But like, it seems pretty insane to do that right now. But there's a coordination problem. Like, you know, Anthropic is worried about if they don't do that, then actually I will do that or opening, I will do that. I'm like, yeah, sure, there is a coordination problem. So let's solve the coordination problem. Like we all recognize that there's this problem. How do we actually solve it? And I think one of the ways we solve it is if we had a lot more transparency into well, who's doing what and the government step in and say like, hey, like you're totally welcome to make amazing products that really advance people's work and lives and help discover cures cancer and make advanced medicines. Like we ******* need that. And we should use all of this intelligence for that. It's a great use. I just don't think we should like try to bootstrap to God like intelligence right away. It just seems like we're not ready for that. we should use all of this intelligence for that. It's a great use. I just don't think we should like try to bootstrap to God like intelligence right away. It just seems like we're not ready for that. And so I'm like having really good monitoring across all this compute infrastructure, knowing where all the chips are, the US and China, knowing where each other's chips are and knowing roughly what we're doing on it. And not being able to set the stage for making deals with each other and saying, hey, like, we both want to go really hard at this kind of stuff. But we see this real danger in these autonomous capabilities that could really undermine human control. Neither of us want that. Let's walk back from the brink here. And you know, it's going to be hard. It's going to, it's going to take our best scientists to come together and figure out how to do this kind of monitoring in ways that are fairly trustless. But I'm like, we have brilliant people working on this. And I think actually technically quite feasible. The biggest difficulty right now is the politics and, and the messaging. But I think it's, that's where to me, that's where most of my hope is, is that we can Orient enough and coordinate enough that we give researchers more time to do the interpretability work we'd need in order to actually know how to trust systems as they're ******* recursive. Myself, improving, I'm like, that's going to be difficult. I think we could probably do it, but we got to grapple with this possibility that we really might need more time. So that's, yeah, that's the main thing I think we got to do. And if we do that, we have a good shot.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence