High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

12 Jul 2026 The Cognitive Revolution Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%

“I think anything that involves slowing down that Frontier has such a low tractability like game theoretically and politically that it kind of this question doesn't matter that much.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
12 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…How much do you think that number is reducible? One story would be we just got to roll the dice at 5% at this point, it is what it is. And another would be we layer on a ton of defence in depth with J space monitoring and natural language auto encoders and constitutional classifiers and probably a few more that we'll come up with or already have and I'm forgetting. And maybe that can take us down to .5%. Like how much kind of marginal impact do you think all these techniques will? Yeah, that's a there. There's a lot of nuances adjacent to this question, but let me start by trying to answer the like the simple, the thing I think you meant to ask and then the higher order considerations. So I think you meant to ask like if we have, as Eliezer calls it, a textbook from the future that explains like what are all the prosaic alignment techniques that actually work? And you applied all of those, like how much of a chance of a misaligned AI would you actually have? I would say 0 like if you actually have, if you, if you actually are kind of mastered the, the theoretical limit of how good prosaic alignment can be, I think you just almost surely in the mathematical sense, like probability one, you'll get an aligned AI, but then there's higher order consideration. So it's like, how long will it take to discover all the prosaic alignment techniques? And, and again, depending on your discount rate, how much you care about, you know, being alive, like how long are you willing to wait? I think you know, again, I do think we're being a little bit reckless. Like if humanity were more coordinated, I think it would make sense to take a pause for about 12 years and accumulate enough prosaic alignment techniques to get it down to like 2% and then and then roll the dice. Like sort of for me, like if I were, you know, in charge of the policy that everyone is going to use to, to, to reason about this, that's sort of the policy that I would prescribe that I think is probably most appropriate. And then there's the question of if we wait 10 years or 12 years or or one year, you know, during that time there are gonna be a lot of techniques that get floated. And some of them like inoculation, prompting might be, in my opinion, not harmful. And so there's like a question of is this kind of gonna wash out? Like if you keep discovering more things? I do think that the things that don't work, you know, there is selection pressure. I think it is self correcting and so more just more prosaic alignment research seems like really good. I do think it on the margin it reduces this. And then there's the there's the question, I guess of yeah, how much, how much is it reducible? Like, yeah, there's a question of feasibility. It's like if you're if you're going to be thinking about theory of change and that's why you're asking this question or that's why you're interested as a listener in this question. I think anything that involves slowing down that Frontier has such a low tractability like game theoretically and politically that it kind of this question doesn't matter that much. So if Eliezer were here, obviously he would disagree with you in terms.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence