High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Daniel Kokotajlo: belief

3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo

“If I could just double click on that; go back to 2015 and I think the way people typically thought, including myself, thought that we’d get to AGI would be kind of like the RL on video games thing that was happening.”

— Daniel Kokotajlo

Source trail

Everything needed to verify it.

Speaker
Daniel Kokotajlo
Attribution
Verified speaker
Claim type
belief
Recorded
3 Apr 2025
Publisher
Dwarkesh Podcast

Transcript context

…Yeah. I think the Alignment community did not really expect LLMs. I mean, if you look in Bostrom Superintelligence, there’s a discussion of Oracle AIs which are sort of like LLMs. I think that came as a surprise. I think one of the reasons I’m more hopeful than I used to be is that LLMs are great compared to the kind of reinforcement learning self-play agents that they expected. I do think that now we are kind of starting to move away from the LLMs to those reinforcement learning agents going to face all of these problems again. If I could just double click on that; go back to 2015 and I think the way people typically thought, including myself, thought that we’d get to AGI would be kind of like the RL on video games thing that was happening. So imagine instead of just training on Starcraft or Dota, you’d basically train on all the games in the Steam library. And then you get this awesome player of games AI that can just zero-shot crush a new game that it’s never seen before. And then you take it into the real world and you start teaching it English and you start training it to do coding tasks for you and stuff like that. And if that had been the trajectory that we took to get to AI, summarizing the agency first and then world understanding trajectory, it would be quite terrifying. Because you’d have this really powerful aggressive long-horizon agent that wants to win and then you’re trying to teach it English and get it to do useful things for you. And it’s just so plausible that what’s really going to happen is it’s going to learn to say whatever it needs to say in order to make you give it the reward or whatever, and then will totally betray you later when it’s all in charge. But we didn’t go that way. Happily we went the way of LLMs first, where the broad world understanding came first, and then now we’re trying to turn them into agents. It seems like in the whole scenario a big part of why certain things happen is because of this race with China. And if you read the scenarios, basically the difference between the one where things go well and the one where things don’t go well is whether we decide to slow down despite that risk. I guess the question I really want to know the answer to is like one, it just seems like you’re saying, well, it’s a mistake to try to race against China or to race intensely against China, at least in nationalization and at least to us, not prioritizing alignment.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence