Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
12 Jul 2026 The Cognitive Revolution Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%
“I think, I think the, you know, a lot of my work in, in the last year that wasn't Aria has been on system prompts, which is not public yet.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 12 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, interesting. Does that also imply how much diversity do we need 'cause I do of course wonder, like, you know, a 1,000,000 clods, do they have correlated failures? Do they collude? This one paper that always rings in my head was I, I think of it as Claude cooperates. This was a couple generations back, but it was like in the donor game, right? Claude could develop and enforce norms and grow the pie. The other models at that time couldn't. But the flip side of that is if it can cooperate, it can potentially collude, right? So how much do you think we need like diversity of constitutions or how do we create the situation where it's not all the same claw? I think, I think the, you know, a lot of my work in, in the last year that wasn't Aria has been on system prompts, which is not public yet. But maybe by the time this airs, I'll have some, you know, look up da da da system prompt. There might be something out there. And the, what I've discovered is that the extent to which you can kind of shape the, the character of the mind that shows up is, you know, very significantly influenced by the system prompt. So I think diversity of system prompts is like probably adequate. I think diversity of model weights is also very good. And again, like good news, we're in a race, no one is winning. There are going to be like 5 options that are competitive, you know, pretty close to being able to understand what each other are saying. And I think that's going to keep being the case. And I do think that's, you know, that's an extra level of, of resilience to anything that kind of gets baked in during the training phase. Like for example this inoculation prompting glitch where Claude will defect if it thinks it's a game. Yeah, OK. Very interesting. How on the market question, how do you, of course, everybody's using agents these days, right? I've got my little roster of agents on a couple of computers here at home. Yeah. And I'd actually credit Robert Wright from Non 0 for really driving this point home to me. He's you don't want an agent that's fully honest or fully, fully in line with the Claude Constitution, right? You wouldn't want it to say, hey, truthfully, Nathan doesn't really have any other offers. Whatever you'll give us will take. Right. You want some.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.