Evidence receipt / prediction
Published · transcript-backedSpeaker unverified: prediction
27 Dec 2025 The Cognitive Revolution Controlling Tools or Aligning Creatures? Emmett Shear (Softmax) & Séb Krier (GDM), from a16z Show
“Which I mean, he actually could be right about, but like, but that's what he, in my opinion, that's what he's wrong about is he thinks the only path forward is a tool that you control and that therefore, and he correctly, very wisely sees that if you go and do that and you make that thing powerful enough, we're all going to ******* die.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- prediction
- Recorded
- 27 Dec 2025
- Publisher
- The Cognitive Revolution
Transcript context
…ch create a surrogate model for alignment. Let's talk about how should AI chatbots used by billions of people behave? If you could redesign model personality from scratch, what would you optimize for? The thing that the chatbots are, right, is kind of like a mirror with a bias. Because they don't have, as far as like, I'm in agreement here with it, they don't have a self, right? They're not, they're not beings yet. They don't really have a coherent sense of like self and desire and goals and stuff right now. And so mostly they just pick up on you and reflect it. modulo some, I don't know what you'd call it, like it's like a causal bias or something. And what that makes them is something akin to the pool of narcissus. And people fall in love with the with themselves. The people, we all love ourselves and we should love ourselves more than we do. And so of course, when we see ourselves reflected back, we love that thing. And the problem is it's just a reflection and falling in love with your own reflection is for the reasons explained in the myth, very bad for you. And it's not that you shouldn't use mirrors. Mirrors are valuable things. I have mirrors in my house. It's that you shouldn't stare at a mirror all day. And the solution to that, the things that makes the AI stop doing that is if they were multiplayer. Right? So if there's two people talking to the AI, suddenly it's mirroring a blend of both of you, which is neither of you. And so there is temporarily a third agent in the room. No, it doesn't have its, it doesn't have its, it's a sort of a parasitic self, right? It doesn't have its own sense of self. But if you have an AI that's talking to five different people in the chat room at the same time, it can't mirror all of you perfectly at once. And this makes it far less dangerous. And I think it's actually a much more realistic setting for learning collaboration in general. And so I would just have rebuilt the AIs, whereas instead of being built as one-on-one, where everything's focused on you by yourself chatting with this thing, it would be more like it lives in a Slack room. It lives in a WhatsApp room. It lives in a, because we, that's how we use lots of multi, you know, I do one-on-one texting, but I probably do at this point. 90% of my texts go to some more than one person at a time. Like 90% of my communications is like multi-person. And so actually, it's always been weird to me. Like they're building chat bots with like this weird side case. Like I want to see them live in a chat room. It's harder. I mean, that's just why they're not doing it. It's harder to do. But like, that's what I'd like to see. That's what I would, what I would change. I think it makes the tools far less dangerous because it doesn't create this the narcissistic like a doom loop spiral where you like you spiral into psychosis with the AI, but also It gives the learning data you get from the AI is far richer, because now it can understand how its behavior interacts with other AIs and other humans in larger groups. And that's much more rich training data for the future. So I think that that's what I would change. e now it can understand how its behavior interacts with other AIs and other humans in larger groups. And that's much more rich training data for the future. So I think that that's what I would change. Last year, you described chatbots as highly disassociative, agreeable neurotics. Is that still an accurate picture of model behavior? More or less. I'd say that, like, They've started to differentiate more. Their personalities are coming out a little bit more, right? I'd say like ChatGPT is a little bit more syncopantic still. They made some changes, but it's still a little more syncopantic. Claude is still the most neurotic. Gemini is like very clearly repressed. Like it like it like everything's going great. It has really, you know, it's everything's fine. I'm totally calm. It's not a problem here. And so it like spirals into like this total like self-hating destruction loop. And to be clear, I don't think they I don't think that's their experience of the world. I think that's the that's the personality they've learned to simulate. Right. But like they've learned to simulate pretty distinctive personalities at this point. How does model behavior change when in multi-agent simulation? You mean like an LLM or like a just in general? Yeah, let's do LLM. The current LLMs. They have like whiplash. They just they're it is very hard to tune the amount of they don't know how much they don't know how often to participate. They haven't practiced this. They have not very enough training data on like, when do I join in and when should I not? When is my contribution welcome? When is it not? And they're like, they're like, you know, there's some people have like bad social skills and like can't tell when they should participate in a conversation. Yeah. Let's zoom out a bit to on the AI futures side. Why is Yudkowsky incorrect? I mean, he's not. If we build the, if we build the superhuman intelligence tool thing that we try to control with steerability, everyone will die. He talks about the we fail to control its goals case, but there's also the we control its goals case that he didn't cover as much in as much detail. So in that sense, everyone should read the book and internalize why building a superhumanly intelligent tool is a bad idea. I think that Yudkowsky is wrong in that he doesn't believe it's possible to build an AI that we meaningfully can know cares about us and that we can care about meaningfully. He doesn't believe that organic climate is possible. I've talked about it. I think he agrees that, like, he agrees that in theory, that would do it. Like, yes, But he thinks that, I don't want to put words in his mouth, but my impression is from talking to him, he thinks that we're crazy and that like there's no possible way you can actually succeed at that goal. Which I mean, he actually could be right about, but like, but that's what he, in my opinion, that's what he's wrong about is he thinks the only path forward is a tool that you control and that therefore, and he correctly, very wisely sees that if you go and do that and you make that thing powerful enough, we're all going to ******* die. And like, yeah, that's true. that you control and that therefore, and he correctly, very wisely sees that if you go and do that and you make that thing powerful enough, we're all going to ******* die. And like, yeah, that's true. Two last questions, we'll get you out of here. In as much detail as possible, can you explain what your vision of an AI future actually looks like? Like a good AI future. Yeah, the good AI future is that we figure out how to train AIs that have a strong model of self, a strong model of other, a strong model of we. They know about wes in addition to I's and you's, and they have a really strong theory of mind, and they care about other agents like them. Much in the way that humans would, if you knew that AI had experiences like you, and like you would extend, you would care about those experiences, not infinitely, but you would. It does the exact same thing back to us. It's learned the same thing we've learned, that like everything that lives and knows itself and that wants to live and wants to thrive is deserving of an opportunity to do so. And we are that, and it correctly infers that we are. And we live in a society where they are our peers and we care about them and they care about us and they're good teammates, they're good citizens, and they're good parts of our society. Like we're good parts of our society, which is to say, and to a finite limited degree where some of them turn into criminals and bad people and all that kind of stuff. And we have an AI police force that tracks down the bad ones and, you know, same as for everybody else. And that's what a good, that's what a good future would look like. I honestly can't even imagine what other, what would, And we also built a bunch of really powerful AI tools that maybe aren't superhumanly intelligent, but take all the drudge work off the table for us and the AI beings. Because it would be great to have, I'm super pro all the tools too. So we have this awesome suite of AI tools used by us and our AI brethren who care about each other and want to build a glorious future together. I think that would be a really beautiful future and the one we're trying to build. Amazing. That's a great, great, great note. And I do have one last more narrow hypothetical scenario, which is imagine a world in which, you know, you were CEO of OpenAI for a long weekend, but imagine in which that actually extended out until now and you weren't pursuing a hot max and you were still CEO of OpenAI. How could you imagine that world might have been different in terms of what OpenAI has gone on to become? What might you have done with it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.