High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Daniel Kokotajlo: belief

3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo

“I think that if the government wades in and tries to make more specific regulations along the lines of what you mentioned, it’s very plausible that it’ll end up backfiring for reasons like what you mentioned.”

— Daniel Kokotajlo

Source trail

Everything needed to verify it.

Speaker
Daniel Kokotajlo
Attribution
Verified speaker
Claim type
belief
Recorded
3 Apr 2025
Publisher
Dwarkesh Podcast

Transcript context

…Let me just add on to that, one of the many other reasons why I worry about nationalization or some kind of public private partnership, or even just very stringent regulation- actually, this is more an argument against very stringent regulation in favor of safety rather than deferring more to the labs on the implementation- is that it just seems like we don’t know what we don’t know about alignment. Every few weeks there’s this new result. OpenAI had this really interesting result recently where they’re like, “hey, they often tell you if they want to hack, in the chain of thought itself. And it’s important that you don’t train against the chain of thought where they tell you they’re going to hack because they’ll still do the hacking if you train against it, they just won’t tell you about it”. You can imagine very naive regulatory responses. It doesn’t just have to be regulations, one might be more optimistic that if it’s an executive order or something, it’ll be more flexible. I just think that relies on a level of goodwill and flexibility on the behalf of our regulator. But suppose there’s some department that says “if you catch your AI saying that they want to take over or do something bad, then you’ll be really heavily punished”. Your immediate response as a lab to just be like, “okay, let’s train them away from saying this”. So you can imagine all kinds of ways in which a top down mandate from the government to the labs of safety would just really backfire, and given how fast things are moving, maybe it makes more sense to leave these kinds of implementation decisions or even high-level strategic decisions around alignment to the labs. Totally, I mean, I also have worried about that exact example. I would summarize the situation as the government lacks the expertise and the companies lack the right incentives. And so it’s a terrible situation. I think that if the government wades in and tries to make more specific regulations along the lines of what you mentioned, it’s very plausible that it’ll end up backfiring for reasons like what you mentioned. On the other hand, if we just trust it to the companies, they’re in a race with each other and they’re full of people who have convinced themselves that this is not a big deal for various reasons and there just is so much incentive pressure for them to win and beat each other and so forth. So even though they have more of the relevant expertise, I also just don’t trust them to do the right things. So Daniel has already said that for this phase we’re not making policy prescriptions. In another phase we may make policy suggestions, and one of the ones that Daniel has talked about that makes a lot of sense to me is to focus on things about transparency. So a regulation saying there has to be whistleblower protection. A big part of our scenario is that a whistleblower comes out and says “the AIs are horribly misaligned and we’re racing ahead anyway”, and then the government pays attention. Or another form of transparency saying that every lab just has to publish their safety case. I’m not as sure about this one because I think they’ll kind of fake it or they’ll publish a made for public consumption safety case that isn’t their real safety case. But at least saying “here is some reason why you should trust us”. And then if all independent researchers say “no, actually you should not trust them”, then I don’t know, they’re embarrassed and maybe they try to do better.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence