Evidence receipt / evaluation
Published · transcript-backedCarl Shulman: evaluation
26 Jun 2023 Dwarkesh Podcast Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future
“If we can probe the thoughts and motivations of an AI and discover wow, actually GPT-6 is planning to takeover the world if it ever gets the chance. That would be an incredibly valuable thing for governments to coordinate around because it would remove a lot of the uncertainty, it would be easier to agree that this was important, to have more give on other dimensions and to have mutual trust that the other side actually also cares about this because you can't always know what another person or another government is thinking but you can see the objective situation in which they're deciding.”
Source trail
Everything needed to verify it.
- Speaker
- Carl Shulman
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Jun 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…To the extent that there is in fact a large risk of AI apocalypse, of all of these governments being overthrown by AI in a way that they don't intend, then it obviously gains from trade and going somewhat slower especially at the end when the danger is highest and the unregulated pace could be truly absurd as we discussed earlier during intelligence explosion. There's no non-competitive reason to try and have that intelligence explosion happen over a few months rather than a couple of years. If you could avert a 10% risk of apocalypse disaster it's just a clear win to take a year or two years or three years instead of a few months to pass through that incredible wave of new technologies without the ability for humans to follow it even well enough to give more proper security supervision, auditing hard power. That's the win. Why might it fail? One important element is just if people don't actually notice a risk that is real so if they just collectively make an error and that does sometimes happen. If it's true this is a probably not-risk then that can be even more difficult. When science pins something down absolutely overwhelmingly then you can get to a situation where most people mostly believe it. Climate change was something that was a subject of scientific study for decades and gradually over time the scientific community converged on a quite firm consensus that human activity releasing carbon dioxide and other greenhouse gases was causing the planet to warm. We've had increasing amounts of action coming out of that. Not as much as would be optimal particularly in the most effective areas like creating renewable energy technology and the like. Overwhelming evidence can overcome differences in people's individual intuitions and priors in many cases. Not perfectly especially when there's political, tribal, financial incentives to look the other way. Like in the United States where you see a significant movement to either deny that climate change is happening or have policy that doesn't take it into account. Even the things that are really strong winds like renewable energy. It's a big problem if as we’re going into this situation when the risk may be very high we don't have a lot of advanced clear warning about the situation. We're much better off if we can resolve uncertainties through experiments where we demonstrate AIs being motivated to reward hack or displaying deceptive appearances of alignment that then break apart when they get the opportunity to do something like get control of their own reward signal. If we could make it be the case in the worlds where the risk is high we know the risk is high, and the worlds where the risk is lower we know the risk is lower then you could expect the government responses will be a lot better. They will correctly note that the gains of cooperation to reduce the risk of accidental catastrophe loom larger relative to the gains of trying to get ahead of one another. onses will be a lot better. They will correctly note that the gains of cooperation to reduce the risk of accidental catastrophe loom larger relative to the gains of trying to get ahead of one another. That's the kind of reason why I'm very enthusiastic about experiments and research that helps us to better evaluate the character of the problem in advance. Any resolution of that uncertainty helps us get better efforts in the possible worlds where it matters the most and hopefully we'll have that and it'll be a much easier epistemic environment. But the environment may not be that easy because deceptive alignment is pretty plausible. The stories we were discussing earlier about misaligned AI involved AI that is motivated to present the appearance of being aligned friendly, honest etc. because that is what we are rewarding, at least in training, and then in training we're unable to easily produce an actual situation where it can do takeover because in that actual situation if it then does it we're in big trouble. We can only try and create illusions or misleading appearances of that or maybe a more local version where the AI can't take over the world but it can seize control of its own reward channel. We do those experiments, we try to develop mind reading for AIs. If we can probe the thoughts and motivations of an AI and discover wow, actually GPT-6 is planning to takeover the world if it ever gets the chance. That would be an incredibly valuable thing for governments to coordinate around because it would remove a lot of the uncertainty, it would be easier to agree that this was important, to have more give on other dimensions and to have mutual trust that the other side actually also cares about this because you can't always know what another person or another government is thinking but you can see the objective situation in which they're deciding. So if there's strong evidence in a world where there is high risk of that risk because we've been able to show actually things like the intentional planning of AIs to do a takeover or being able to show model situations on a smaller scale of that I mean not only are we more motivated to prevent it but we update to think the other side is more likely to cooperate with us and so it's doubly beneficial. Famously in the game theory of war, war is most likely when one side thinks the other is bluffing but the other side is being serious or when there's that kind of uncertainty. If you can prove the AI is misaligned you don't think they're bluffing about not wanting to have an AI takeover, right? You can be pretty sure that they don't want to die from AI.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.