Evidence receipt / belief
Published · transcript-backedScott Alexander: belief
3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo
“I think with race, it’s very easy to know whether you’re black or white, and so there have been many cases of one race kind of conspiring against another for a long time, like apartheid or any of the racial genocides that have happened.”
Source trail
Everything needed to verify it.
- Speaker
- Scott Alexander
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Apr 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah. So it seems like this community is very interested in solving this problem at a technical level of making sure AIs don’t lie to us, or maybe they lie to us in the scenarios exactly where we would want them to lie to us or something. Whereas as you were saying, humans have these exact same problems. They reward hack, they are unreliable, they obviously do cheat and lie. And the way we’ve solved it with humans is just checks and balances, decentralization. You could lie to your boss and keep lying to your boss, but over time it’s just not going to work out with you- or you become president or something, one or the other. So if you believe in this extremely fast take off, if a lab is one month ahead, then that’s the end game and this thing takes over. But even then- I know I’m combining so many different topics- even then, there’s been a lot of theories in history which have had this idea of “some class is going to get together and unite against the other class”. And in retrospect, whether it’s the Marxist, whether it’s people who have some gender theory or something, like the proletariat will unite or the females will unite or something, they just tend to think that certain agents have shared interests and will act as a result of the shared interest in a way that we don’t actually see in the real world. And in retrospect, it’s like, “wait, why would all the proletariat like…” So why think that this lab will have these AIs who are… there’s a million parallel copies and they all unite to secretly conspire against the rest of human civilization in a way that, even if they are deceitful in some situations. I kind of want to call you out on the claim that groups of humans don’t plot against other groups of humans. I do think we are all descended from the groups of humans who successfully exterminated the other groups of humans, most of whom throughout history have been wiped out. I think even with questions of class, race, gender, things like that, there are many examples of the working class rising up and killing everybody else. And if you look at why this happens, why this doesn’t happen, it tends to happen in cases where one group has an overwhelming advantage. This is relatively easy for them. You tend to get more of a diffusion of power democracy where there are many different groups and none of them can really act on their own. And so they all have to form a coalition with each other. There’s also cases where it’s very obvious who’s part of what group. So for example, with class, it’s hard to tell whether the middle class should support the working class versus the aristocrats. I think with race, it’s very easy to know whether you’re black or white, and so there have been many cases of one race kind of conspiring against another for a long time, like apartheid or any of the racial genocides that have happened. I do think that AI is going to be more similar to the cases where, number one, there’s a giant power imbalance, and number two, they are just extremely distinct groups that may have different interests. I think I’d also mention the homogeneity point. Any group of humans, even if they’re all exactly the same race and gender, is going to be much more diverse than the army of AIs in the data center, because they’ll mostly be literal copies of each other. And I think that goes for a lot. Another thing I was going to mention is that our scenario doesn’t really explore this. I think in our scenario, they’re more like a monolith. But historically, a lot of crazy conquests happened from groups that were not at all monoliths. And I’ve been heavily influenced by reading the history of the conquistadors, which you may know about. But did you know that when Cortez took over Mexico, he had to pause halfway through, go back to the coast, and fight off a larger Spanish expedition that was sent to arrest him? So the Spanish were fighting each other in the middle of the conquest of Mexico. Similarly, in the conquest of Peru, Pizarro was replicating Cortez’s strategy, which, by the way, was “go get a meeting with the emperor and then kidnap the emperor and force him at sword point to say that actually everything’s fine and that everyone should listen to your orders”. That was Cortez’s strategy, and it actually worked. And then Pizarro did the same thing, and it worked with the Inca. But also with Pizarro, his group ended up getting into a civil war in the middle of this whole thing. And one of the most important battles of this whole campaign was between two Spanish forces fighting it out in front of the capital city of the Incas. And more generally, the history of European colonialism is like this, where the Europeans were fighting each other intensely the entire time, both on the small scale within individual groups, and then also at the large scale between countries. And yet nevertheless they were able to carve up the world and take over. And so I do think this is not what we explore in the scenario, but I think it’s entirely plausible that even if the AIs within an individual company are in different factions, they might nevertheless overall end up quite poorly for humans.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.