High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Daniel Kokotajlo: uncertainty

3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo

“We don’t know what actual goals will end up inside the AIs and what the sort of internal structure of that will be like, what goals will be instrumental versus terminal.”

— Daniel Kokotajlo

Source trail

Everything needed to verify it.

Speaker
Daniel Kokotajlo
Attribution
Verified speaker
Claim type
uncertainty
Recorded
3 Apr 2025
Publisher
Dwarkesh Podcast

Transcript context

…succeed, really likes making money, really likes the thrill of successful tasks. They’re also being regulated and they’re like, “yeah, I guess I’ll follow the regulation, I don’t want to go to jail”. But it is not robustly, deeply aligned to, “yes, I love regulations, my deepest drive is to follow all of the regulations in my industry”. So we think that an AI like that, as time goes on and as this recursive self improvement process goes on, will kind of get worse rather than better. It will move from kind of this vague superposition of “well, I want to succeed, I also want to follow things” to being smart enough to genuinely understand its goal system and being like, “my goal is success, I have to pretend to want to do all of these moral things while the humans are watching me”. That’s what happens in our story. And then at the very end, the AIs reach a point where the humans are pushing them to have clearer and better goals because that’s what makes the AIs more effective. And they eventually clarify their goals so much that they just say, “yes, we want task success. We’re going to pretend to do all these things well while the humans are watching us”. And then they outgrow the humans and then there’s disaster. To be clear, we’re very uncertain about all of this. So we have a supplementary page on our scenario that goes over different hypotheses for what types of goals AIs might develop in training processes similar to the ones that we are depicting, where you have these lots of agency training, you’re making these AI agents that autonomously operate, doing all this ML R&D, and then you’re rewarding them based on what appears to be successful. And you’re also slapping on some sort of alignment training as well. We don’t know what actual goals will end up inside the AIs and what the sort of internal structure of that will be like, what goals will be instrumental versus terminal. We have a couple different hypotheses and we picked one for purposes of telling the story. I’m happy to go into more detail if you want, about the mechanistic details of the particular hypothesis we picked or the different alternative hypotheses that we didn’t depict in the story that also seem plausible to us. Yeah, we don’t know how this will work at the limit of all these different training methods, but we’re also not completely making this up. We have seen a lot of these failure modes in the AI agents that exist already.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence