High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Scott Alexander: prediction

3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo

“We’re telling it “succeed on tasks, which is going to make you a power seeker, but also don’t seek power in these particular ways”. And in our scenario, we predict that this doesn’t work and that the AI learns to seek power and then hide it.”

— Scott Alexander

Source trail

Everything needed to verify it.

Speaker
Scott Alexander
Attribution
Verified speaker
Claim type
prediction
Recorded
3 Apr 2025
Publisher
Dwarkesh Podcast

Transcript context

…Well, that’s what it is. And part of the reason for that is just that I feel like a bunch of stuff has to go right. I feel like we can’t just unilaterally slow down and have China go take the lead. That also is a terrible future. But we can’t also completely race, because for the reasons I mentioned previously about alignment, I think that if we just go all out on racing, we’re going to lose control of our AIs, right? And so we have to somehow thread this needle of pivoting and doing more alignment research and stuff, but not too much that helps China win. And that’s all just for the alignment stuff. But then there’s the concentration of power stuff where somehow in the middle of doing all of that, the powerful people who are involved need to somehow negotiate a truce between themselves to share power and then ideally spread that power out amongst the government and get the legislative branch involved. Somehow that has to happen too, otherwise you end up with this horrifying dictatorship or oligarchy. It feels like all that stuff has to go right and we depict it all going mostly right in one ending of our story. But yeah, it’s kind of rough. So I am the writer and the celebrity spokesperson for this scenario. I am the only person on the team who is not a genius forecaster. And maybe related to that, my p(doom) is the lowest of anyone on the team. I’m more like 20%. First of all, people are going to freak out when I say this. I’m not completely convinced that we don’t get something like alignment by default. I think that we’re doing this bizarre and unfortunate thing of training the AI in multiple different directions simultaneously. We’re telling it “succeed on tasks, which is going to make you a power seeker, but also don’t seek power in these particular ways”. And in our scenario, we predict that this doesn’t work and that the AI learns to seek power and then hide it. I am pretty agnostic as to exactly what happens. Maybe it just learns both of these things in the right combination, I know there are many people who say that’s very unlikely. I haven’t yet had the discussion where that worldview makes it into my head consistently. And then I also think we’re going to be involved in this race against time. We’re going to be asking the AIs to solve alignment for us. The AIs are going to be solving alignment because even if they’re misaligned, they want to align their successors. So they’re going to be working on that. And we have these two competing curves. Can we get the AI to give us a solution for alignment before our control of the AI fails so completely that they’re either going to hide their solution from us, or deceive us, or screw us over in some other way? That’s another thing where I don’t feel like I have any idea of the shape of those curves. I’m sure if it were Daniel or Eli, they would have already made five supplements on this. But for me, I’m just kind of agnostic as to whether we get to that alignment solution, which in our scenario, I think we focus on mechanistic interpretability. Once we can really understand the weights of an AI on a deep level, then we have a lot of alignment techniques open up to us. I don’t really have a great sense of whether we get that before or after the AI has become completely uncontrollable. And a big part of that relies on the things we’re talking about. How smart are the labs? How carefully do they work on controlling the AI? How long do they spend making sure the AI is actually under control and the alignment plan they gave us is actually correct, rather than something they’re trying to use to deceive us? All of those things I’m completely agnostic on, but that leaves like a pretty big chunk of probability space where we just do okay. And I admit that my p(doom) is literally just p(doom) and not p(doom or oligarchy). So that 80% of scenarios where we survive contains a lot of really bad things that I’m not happy about. But I do think that we have a pretty good chance of surviving. Let’s talk about geopolitics next. So describe to me how you foresee the relationship between the government and the AI labs to proceed, how you expect that relationship in China to proceed, and how you expect the relationship between the US and China to proceed. Okay, three simple questions. Yes, no, yes, no, yes, no.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence