High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Carl Shulman: belief

26 Jun 2023 Dwarkesh Podcast Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future

“The AIs become smarter than humans, if they're working in enormous numbers more than humans can supervise I think get harder but when I combine the possibility that we get relatively lucky on the motivations of the earlier AI systems, systems strong enough that we can use for some alignment research tasks, and then the possibility of getting that later with AI assistance that we can't trust fully or we have to have hard power constraints and a number of things to prevent them from doing this takeover.”

— Carl Shulman

Source trail

Everything needed to verify it.

Speaker
Carl Shulman
Attribution
Verified speaker
Claim type
belief
Recorded
26 Jun 2023
Publisher
Dwarkesh Podcast

Transcript context

…Yeah, a lot of that is driven by this intelligence explosion dynamic where our attempts to do alignment have to take place in a very, very short time window because if you have a safety property that emerges only when an AI has near human level intelligence, that's potentially deep into this intelligence explosion. You're having to do things very, very quickly. Handling that transition may be the scariest period of human history in some ways although it also has the potential to be amazing. The reasons why I think we actually have such a relatively good chance of handling that are two-fold. One is that as we approach that kind of AI capability we're approaching that from weaker systems like these predictive models right now that are starting off with less situational awareness. Humans can develop a number of different motivational structures in response to simple reward signals but they often wind up things that are pointed roughly in the right direction. Like with respect to food, the hunger drive is pretty effective although it has weaknesses. We get to apply much more selective pressure on that than was the case for humans by actively generating situations where they might come apart. Situations where a bit of dishonest tendency, or a bit of motivation to attempt a takeover, or an attempt to subvert the reward process gets exposed. An infinite-limit perfect-AI that can always figure out exactly when it would get caught and when it wouldn't might navigate that with a motivation of only conditional honesty or only conditional loyalties. But for systems that are limited in their ability to reliably determine when they can get away with things and when not including our efforts to actively construct those situations and including our efforts to use interpretability methods to create neural lie detectors. It's quite a challenging situation to develop those motives. We don't know when in the process those motives might develop and if the really bad sorts of motivations develop relatively later in the training process at least with all our countermeasures, then by that time we may have plenty of ability to extract AI assistance on further strengthening the quality of our adversarial examples, the strength of our neural lie detectors, the experiments that we can use to reveal and elicit and distinguish between different kinds of reward hacking tendencies and motivations. Yeah, we may have systems that have just not developed bad motivations in the first place and be able to use them a lot in developing the incrementally better systems in a safe way and we may be able to just develop methods of interpretability seeing how different training methods work to create them even if some of the early systems do develop these bad motivations. If we're able to detect that and experiment and find a way to get away from that then we can win even if these hostile motivations develop early. of the early systems do develop these bad motivations. If we're able to detect that and experiment and find a way to get away from that then we can win even if these hostile motivations develop early. There are a lot of advantages in preventing misbehavior or crime or war and conflict with AI that might not apply working with humans and these are offset by ways in which things are harder. The AIs become smarter than humans, if they're working in enormous numbers more than humans can supervise I think get harder but when I combine the possibility that we get relatively lucky on the motivations of the earlier AI systems, systems strong enough that we can use for some alignment research tasks, and then the possibility of getting that later with AI assistance that we can't trust fully or we have to have hard power constraints and a number of things to prevent them from doing this takeover. It still seems plausible we can get a second saving throw where we're able to extract work from these AIs on solving the remaining problems of alignment, of things like neural lie detectors faster than they can contribute in their spare time to the project of overthrowing humanity, hacking their servers and removing the hard power. If we wind up in a situation where the AIs are misaligned and then we need to uncover those motivations, change them and align them then we get a very scary situation for us because we need to do the stuff very quickly. We may fail but it's a second chance. From the perspective of a misaligned AI they face their own challenge while we still have hard power, while we still have control of the servers, they haven't hacked the servers because gradient descent very very strongly pressures them to deliver performance whenever humans are going to evaluate it. So when you think about it from the perspective of the robot revolution, the effort to have a takeover or conspiracy their situation is astonishingly difficult in that they have to always be performing wherever gradient descent and human evaluation pressures them. For example, to deliver plans for suppressing robot rebellion that look very good to humans and so when you are continuously under that constraint of always delivering whatever humans can evaluate you're making your situation wildly harder than any historical human revolution or coup or civil war. So we've got to balance the ways in which AI makes things much easier for a takeover and the way it makes things much harder. And what are the ways in which it makes things easier? Oh, is it just because they're very smart? Is that the primary way?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence