Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
24 May 2026 The Cognitive Revolution All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
“We, I think there's this big dream, which I'm excited about of like a is taking search costs super low, kind of facilitating all these transactions that previously couldn't have happened because the transaction costs were too high to facilitate.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 24 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Very little. I do think we should note that this is very interesting. I will say I have been surprised at how good Claude is at saying moral things. I'm like Claude as if you ask Claude for ethics situation, Claude can give you pretty damn good advice. It's really impressive. And I think that that means something and is significant makes me like marginally more optimistic. Unfortunately, the reason why it doesn't go very deep to me is that I just think there's a huge difference between training something to say good things in training something to act morally and especially have moral motivations, underlying motivations. I think that even though Claude says very moral things and can give you very moral advice, Claude is still pretty amoral in some sense. And what I would say is the task of training a model to say very moral things is hard, but less hard than the task of getting an A model to solve a totally novel math problem or something, or like, figure out a new material and test the material. And I think this really matters because in some sense, you have this weird. I'm like, Nathan, if you were I asked you for advice on a bunch of moral questions in my life, OK, I'm having this interpersonal conflict. What should I do? And you gave me the advice of Claude. I'd be like, Nathan's a really good guy and at the same time, if you lied as much as Claude lies, I'd be like Nathan, it's a total immoral. Like he's terrible. He you can't trust him. He's not a trustworthy guy. This is very confusing. Like it wouldn't make any sense. Also, I'd have to question, I'd be like, am I OK, Nathan be this immoral, this moral and this immoral at the same time. That's like a very unusual thing in humans. It's not unusual in models. In fact, it's basically the default in models. Like if you go ask Grok moral questions, Grok's pretty moral too. And so is Chad TPT, and so is Gemini. And also these guys lie all the time, and they cheat all the time. And this says something interesting about the ways in which whether the model says good things and does good things is less connected than it is in humans. And unfortunately, I think this means we have to be very careful. And I think, sorry, I guess in my take is that I think a lot of people are misled by this and hear the model saying moral things and assume that that means that the model has good motivations. And I just don't think that's really the case. Yeah, it's certainly not something we should be taking for granted. I can say that with 100% confidence. I've talked about this many times, but a formative experience for me was doing the GPT 4 Red team and using the helpful only model and just realizing how vast the space of AI mines really is and how easy it is for them to end up in a state that really violates our intuition for how people are going to be. As you said, much more correlated along different dimensions than minds in general have to be or that a is have to be. And so that's definitely something we should be keeping in mind a lot. I think another big interesting thing that's coming coming up on the alignment frontier is multi agent competitive world where mostly so far we've trained things to handle one thread, pursue 1 task for one user without too much in the way of dynamics. I'm sure you've followed and in labs work with their vending bench and things like that. It's been interesting to see recently that the most recent claws they've described as being ruthless. I guess it's kind of there's a comment here and then there's I think a question as well. The observation is that's kind of a Yikes. And I don't know, obviously I don't know all that's going on in quad training. But it sure seems like we are entering A regime where one of the very natural next things to do is going to be to train agents in competitive environments where they're supposed to make money or they're supposed to negotiate or they're supposed to represent interests in a world where other agents or entities have other interests. That seems like it's going to be a big Yikes because that the world itself just rewards deception. We see deception in nature all over the place. And there's a very fundamental reason for that, which is you can win by deceiving the other agents in your environment. So I'm really on the lookout right now for how are companies going to handle that if they want their AIS to be able to go out? And by the way, the economy is an adversarial environment, right? Like if you are naive and you go out into the world of suppliers and negotiations or whatever, and you take everything at face value and you don't try to push back a little bit. Or if, if you don't have some separation between your initial offer and your bottom line reservation price, then you're going to be the sucker that's taken advantage of. And nobody's going to want to use that AI to go out and do these sorts of things, right? We, I think there's this big dream, which I'm excited about of like a is taking search costs super low, kind of facilitating all these transactions that previously couldn't have happened because the transaction costs were too high to facilitate. But to do that well, they are going to have to have certain amount of at least minimal deception. And it seems like we're we're already kind of seeing it arise through sort of whatever prior and accident and there's now seemingly very already too soon. certain amount of at least minimal deception. And it seems like we're we're already kind of seeing it arise through sort of whatever prior and accident and there's now seemingly very already too soon. There's like a very direct incentive to dial that up and I don't know how we're going to figure out how to balance that. So a comment that is very.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.