Evidence receipt / belief
Published · transcript-backedDaniel Kokotajlo: belief
3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo
“I think we’re substantially more likely to make things get the technical stuff right if it’s something like that.”
Source trail
Everything needed to verify it.
- Speaker
- Daniel Kokotajlo
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Apr 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…I think traditionally in the AI safety community there’s been this idea which I myself used to believe, that it’s an incredibly high priority to basically have way better information security. And if you’re going to be trying to build AGI, you should not be publishing your research, because that helps other less responsible actors build AGI. And the whole game plan is for a responsible actor to get to AGI first and then stop and burn down their lead time over everybody else and spend that lead on making it safe, and then proceed. And so if you’re publishing all your research, then there’s less lead time because your competitors are going to be close behind you. And other reasons too, but that’s one reason why I think historically people such as myself have been pro-secrecy. Another reason, of course, is obviously you don’t want rivals to be stealing your stuff. But I think that I’ve now become somewhat disillusioned and think that even if we do have a three-month lead, a six-month lead, between the leading US project and any serious competitor, it’s not at all foregone conclusion that they will burn that lead for good purposes, either for safety or for constitutional power stuff. I think the default outcome is that they just smoothly continue on without any serious refocusing. And part of why I think this is because this is what a lot of the people at the company seem to be planning and saying they’re going to do. A lot of them are basically like “the AIs are just going to be misaligned by then. They seem pretty good right now. Oh yeah, sure, there were a few of those issues that various people have found, but we’re ironing them out. It’s no big deal”. That’s what a huge amount of these people think. And then a bunch of other people think, even though they are more concerned about misalignment, they’ll figure it out as they go along and there won’t need to be any substantial slowdown. Basically, I’ve become more disillusioned that they’ll actually use that lead in any sort of reasonable, appropriate way. And then I think that separately, there’s just a lot of intellectual progress that has to happen for the alignment problem to be more solved than it currently is now. I think that currently there’s various alignment teams at various companies that aren’t talking that much with each other and sharing their results. They’re doing a little bit of sharing and a little bit of publishing like we’re seeing, but not as much as they could. And then there’s a bunch of smart people in academia that are basically not activated because they don’t take all this stuff seriously yet, and they’re not really waking up to superintelligence yet. And what I’m hoping will happen is that this situation will get better as time goes on. What I would like to see is society as a whole starting to freak out as the trend lines start upwards and things get automated and you have these fully autonomous agents and they start using neuralese and hive mind. ke to see is society as a whole starting to freak out as the trend lines start upwards and things get automated and you have these fully autonomous agents and they start using neuralese and hive mind. As all that exciting stuff starts happening in the data centers, I would like it to be the case that the public is following along and then getting activated and all of these other researchers are reading the safety case and critiquing it and doing little ML experiments on their own tiny compute clusters to examine some of the assumptions in the safety case and so forth. Basically, one way of summarizing it is that currently there’s going to be 10 alignment experts in whatever inner silo of whatever company is in the lead. And the technical issue of making sure that AIs are actually aligned is going to fall roughly to them. But what I would like to be is a situation where it’s more like 100 or 500 alignment experts spread out over different companies and in nonprofits that are sort of all communicating with each other and working on this together. I think we’re substantially more likely to make things get the technical stuff right if it’s something like that. Let me just add on to that, one of the many other reasons why I worry about nationalization or some kind of public private partnership, or even just very stringent regulation- actually, this is more an argument against very stringent regulation in favor of safety rather than deferring more to the labs on the implementation- is that it just seems like we don’t know what we don’t know about alignment. Every few weeks there’s this new result. OpenAI had this really interesting result recently where they’re like, “hey, they often tell you if they want to hack, in the chain of thought itself. And it’s important that you don’t train against the chain of thought where they tell you they’re going to hack because they’ll still do the hacking if you train against it, they just won’t tell you about it”. You can imagine very naive regulatory responses. It doesn’t just have to be regulations, one might be more optimistic that if it’s an executive order or something, it’ll be more flexible. I just think that relies on a level of goodwill and flexibility on the behalf of our regulator. But suppose there’s some department that says “if you catch your AI saying that they want to take over or do something bad, then you’ll be really heavily punished”. Your immediate response as a lab to just be like, “okay, let’s train them away from saying this”. So you can imagine all kinds of ways in which a top down mandate from the government to the labs of safety would just really backfire, and given how fast things are moving, maybe it makes more sense to leave these kinds of implementation decisions or even high-level strategic decisions around alignment to the labs.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.