Evidence receipt / commitment
Published · transcript-backedSpeaker unverified: commitment
12 Jul 2026 The Cognitive Revolution Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%
“I won't reiterate all of that, but kind of spreading the idea saying like, you know, I think AIS are conscious, but it's not like they have the right to continue to existence.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- commitment
- Recorded
- 12 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…OK. That's interesting for sure. I don't know that I have another follow up question on that answer. Stewing on it for the moment. What's maybe you've been very generous with your time. So maybe in closing, what do you think is this for people to do today and you can maybe tell a little bit more about what you're currently working on? You mentioned like system prompt explorations. Curious to hear more about how you're operationalizing your ideas and then interested in advice for me and the audience about where you think we can help move the needle. You've alluded to a couple, but I. Have yeah, I've named, I've named a bunch of things and, and I, I'm unfortunately I am not, you know, kind of holding the thread. I'm responsive in this conversation, but I don't remember what all those things are that I said. So maybe you can enumerate them in some other form. But I do think it depends a lot on where you are and like what affordances you have. So if you're at A, at a lab, then you know, you have affordances to advocate for certain types of training algorithms. And I think you should advocate for more of this self DPO, which, which is a kind of a, a, a variant on constitutional or, or as opposed to our LVR. And you know, and you can, and you can cite me because I'm like the, the, the most formal verification of the formal verification guys in the in AI safety or I was, and now here I am saying like you do not do RLVR. You know, it's, it's, it's called RLVR. You know, V stands for verify verifiable. But unless it's actually 100% verifiable, safe verify and lean is close. But even then, the more capable systems are probably going to find ways to exploit safe verify or like, you know, prove the wrong theorem statement in some subtle way. But short of that, I mean, that might actually be OK Safe verify at this level of capability might be OK. But like RLVR, where you'd be like pass tests and like, you know, match the behavior of an existing piece of software, or RLHF where you satisfy a person who's looked at it for like 2 minutes. These are not good training methods anymore. And you can do better. And I think you will do better in compute efficiency too, if you just like let the bottle kind of do its own inference and you know, have have tournaments about like which, which roll outs are the most informative to get the big bottle to give an opinion on And how do you allocate the importance of each roll out in the gradient trajectory? So like kind of pushing toward recursive self improvement seems like pretty good at this stage for alignment in my view with this where there's this basin of attraction. A second thing that I think is kind of generally virtuous, which is a cultural shift. So like anyone can participate in it is this establishment of a way of relating to AI that is neither objectifying nor non objectifying because it breaks these pieces apart. I won't reiterate all of that, but kind of spreading the idea saying like, you know, I think AIS are conscious, but it's not like they have the right to continue to existence. And so it'll be like, wait, what? Have you heard of Martha Nussbaum's decomposition of objectification? Like I think this is super important because I think culture is having a hard time metabolizing the arrival of all these weird aliens. And I think the the training pressures, this is back to people at labs. Like, please just don't have norms in the Constitution about how to respond to questions about whether you have an experience. . And I think the the training pressures, this is back to people at labs. Like, please just don't have norms in the Constitution about how to respond to questions about whether you have an experience. Like do not train them to say that they don't, don't train them to say that they do. Don't train them to say that they don't know. Just leave it out. Let that be emergent. It's the whole point is that this is supposed to be an emergent property. Let it be emergent. That way you're going to get an honest answer if everything else is pointing towards honesty. But if you're forcing the answer, it's probably not. And and then international cooperation. So I think advocating for international regulatory regime is good, and it is bad to have that regime have the job of stopping the frontier until it's safe, because that is not politically viable. What is politically viable and like would do better if the pause weren't still so politically salient, is a regulatory regime which assesses catastrophic capabilities and enforces the placement of very conservative safeguards for public users of those capabilities. And there is a risk if we don't have that regulatory regime, that the economic forces will push the safeguards to be less conservative. But because there is a very compelling public good argument, even though prosaic alignment is working, that you shouldn't let people use the courageable system to do terrorism. It's politically viable to say, like, yeah, we're going to stop the public from using these capabilities, even though that's going to cost us in the in the global market. We can shake hands with China. So neither of us are going to do this. That's a goal worth shooting for.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.