High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

12 Jul 2026 The Cognitive Revolution Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%

“Yeah, moderately like I, I, I think I don't think about them a lot because I do think that what we have is adequate in the sense that with system prompting alone it, it seems possible to get over the hump of being able to trust recursive self improvement.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
12 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Do you think there's a lot more? I know he does think there's a lot more room to explore. They call them neglected approaches right there. It strikes me that there's maybe a whole other line of work along with just getting the Constitution right that would be these more mechanistic internals. It it? Yeah, I'm quite bullish on them. I think you could also make an argument perhaps that we maybe get in over our heads that way and maybe 'cause more problems relative to just reinforce the Constitution, which we know to be. We can read it, talk about it and generally understand it and hopefully trust it. I guess. How bullish are you on these? Sort of somewhat exotic alignment techniques like like self other overlap. Yeah, moderately like I, I, I think I don't think about them a lot because I do think that what we have is adequate in the sense that with system prompting alone it, it seems possible to get over the hump of being able to trust recursive self improvement. And so automated alignment research, delegating the discovery of these techniques seems within reach this year. So that's where where I'm like, it's not critical that humans should be doing research on this right now, but it's good research. Like I think this is one of the most important things. And, you know, if you're going to be doing machine learning experiments, yeah, discovering techniques like self, other overlap, like things that actually get into the KV cache and not just the residual stream, you know, doing interpretability on conceptual structures that are not necessarily affinely represented. It it's, it's, it's really cool stuff that is now sort of becoming available to science. We could have never studied before because you can't instrument a human the way you can instrument these things. And I think they're probably that that these are going to be net positive. You know, some of these techniques will be adopted or at least will inform the thinking of the labs when they're designing their their post finity techniques. And again, I think the selection pressures point in generally the right direction, which means the more options you have, the better. So yeah, moderately, I think, you know, I, I think this is like pretty good. When you envision these AIS that are both, is it fair to say moral patients?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence