Evidence receipt / uncertainty
Published · transcript-backedDwarkesh Patel: uncertainty
3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo
“Let me just add on to that, one of the many other reasons why I worry about nationalization or some kind of public private partnership, or even just very stringent regulation- actually, this is more an argument against very stringent regulation in favor of safety rather than deferring more to the labs on the implementation- is that it just seems like we don’t know what we don’t know about alignment.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 3 Apr 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…ke to see is society as a whole starting to freak out as the trend lines start upwards and things get automated and you have these fully autonomous agents and they start using neuralese and hive mind. As all that exciting stuff starts happening in the data centers, I would like it to be the case that the public is following along and then getting activated and all of these other researchers are reading the safety case and critiquing it and doing little ML experiments on their own tiny compute clusters to examine some of the assumptions in the safety case and so forth. Basically, one way of summarizing it is that currently there’s going to be 10 alignment experts in whatever inner silo of whatever company is in the lead. And the technical issue of making sure that AIs are actually aligned is going to fall roughly to them. But what I would like to be is a situation where it’s more like 100 or 500 alignment experts spread out over different companies and in nonprofits that are sort of all communicating with each other and working on this together. I think we’re substantially more likely to make things get the technical stuff right if it’s something like that. Let me just add on to that, one of the many other reasons why I worry about nationalization or some kind of public private partnership, or even just very stringent regulation- actually, this is more an argument against very stringent regulation in favor of safety rather than deferring more to the labs on the implementation- is that it just seems like we don’t know what we don’t know about alignment. Every few weeks there’s this new result. OpenAI had this really interesting result recently where they’re like, “hey, they often tell you if they want to hack, in the chain of thought itself. And it’s important that you don’t train against the chain of thought where they tell you they’re going to hack because they’ll still do the hacking if you train against it, they just won’t tell you about it”. You can imagine very naive regulatory responses. It doesn’t just have to be regulations, one might be more optimistic that if it’s an executive order or something, it’ll be more flexible. I just think that relies on a level of goodwill and flexibility on the behalf of our regulator. But suppose there’s some department that says “if you catch your AI saying that they want to take over or do something bad, then you’ll be really heavily punished”. Your immediate response as a lab to just be like, “okay, let’s train them away from saying this”. So you can imagine all kinds of ways in which a top down mandate from the government to the labs of safety would just really backfire, and given how fast things are moving, maybe it makes more sense to leave these kinds of implementation decisions or even high-level strategic decisions around alignment to the labs. Totally, I mean, I also have worried about that exact example. I would summarize the situation as the government lacks the expertise and the companies lack the right incentives. And so it’s a terrible situation. I think that if the government wades in and tries to make more specific regulations along the lines of what you mentioned, it’s very plausible that it’ll end up backfiring for reasons like what you mentioned. On the other hand, if we just trust it to the companies, they’re in a race with each other and they’re full of people who have convinced themselves that this is not a big deal for various reasons and there just is so much incentive pressure for them to win and beat each other and so forth. So even though they have more of the relevant expertise, I also just don’t trust them to do the right things.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.