High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Carl Shulman: prediction

26 Jun 2023 Dwarkesh Podcast Carl Shulman (Pt 2) — AI Takeover, bio & cyber attacks, detecting deception, & humanity's far future

“If in the future it's as easy to set the actual underlying motivations of AI as it is right now to set the behavior that they display then it means you could have AI's created with almost whatever motivation people wish and that could really drastically change political affairs because the ability to decide and determine the loyalties of the humans or AIs and robots that hold the guns, that hold together society, that ultimately back it against violent overthrow and such.”

— Carl Shulman

Source trail

Everything needed to verify it.

Speaker
Carl Shulman
Attribution
Verified speaker
Claim type
prediction
Recorded
26 Jun 2023
Publisher
Dwarkesh Podcast

Transcript context

…I think that's a really good lead-in into the topic of lock-in. You just mentioned how there can be these kinds of coups if a large portion of the population is unsatisfied with the regime, why might this not be the case with superhuman intelligences in the far future? I also said it specifically with respect to things like security forces and the sources of hard power. In human affairs there are governments that are vigorously supported by a minority of the population, some narrow electorate that gets treated especially well by the government while being unpopular with most of the people under their rule. We see a lot of examples of that and sometimes that can escalate to civil war when the means of power become more equally distributed or there's a foreign assistance provided to the people who are on the losing end of that system. Going forward, I don't expect that definition to change. I think it will still be the case that a system that those who hold the guns and equivalent are opposed to is in a very difficult position. However AI could change things pretty dramatically in terms of how security forces and police and administrators and legal systems are motivated. Right now we see with GPT-3 or GPT-4 that you can get them to change their behavior on a dime. So there was someone who made a right-wing GPT because they noticed that on political compass questionnaires the baseline GPT-4 tended to give progressive San Francisco type of answers which is in line with the people who are providing reinforcement learning data and to some extent reflecting like the character of the internet. So they did a little bit of fine-tuning with some conservative data and then they were able to reverse the political biases of the system. If you take the initial helpfulness-only trained models for some of these over, I think there's anthropic and OpenAI have published both some information about the models trained only to do what users say and not trained to follow ethical rules, and those models will behaviorally eagerly display their willingness to help design bombs or bioweapons or kill people or steal or commit all sorts of atrocities. If in the future it's as easy to set the actual underlying motivations of AI as it is right now to set the behavior that they display then it means you could have AI's created with almost whatever motivation people wish and that could really drastically change political affairs because the ability to decide and determine the loyalties of the humans or AIs and robots that hold the guns, that hold together society, that ultimately back it against violent overthrow and such. It's potentially a revolution in how societies work compared to the historical situation where security forces had to be drawn from some broader populations, offered incentives, and then the ongoing stability of the regime was dependent on whether they remained bought in to the system. This is slightly off topic but one thing I'm curious about is what does the median far future outcome of AI look like? Do we get something that, when it has colonized the galaxy, is interested in diverse ideas and beautiful projects or do we get something that looks more like a paper-clip maximizer? Is there some reason to expect one or the other? I guess what I'm asking is, there's some potential value that is realizable within the matter of this galaxy. What does the median outcome look like compared to how good things could be?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence