High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Dwarkesh Podcast / episode intelligence

Paul Christiano — Preventing an AI takeover

31 Oct 2023 53 published claims 2 attributable people

Speakers in the public record

Claim mix

belief 29uncertainty 9prediction 6evaluation 5commitment 3preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

53 published records

01 / prediction

Yeah, I mean, I think amongst people you've interviewed, maybe that's like on the long end thinking it would take like a couple of years. And it depends a little bit what you mean by I think literally all human cognitive labor is probably like more like weeks or months or something like that.

“Yeah, I mean, I think amongst people you've interviewed, maybe that's like on the long end thinking it would take like a couple of years. And it depends a little bit what you mean by I think literally all human cognitive labor is probably like more like weeks or months or something like that.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

05 / belief

You think you can understand or control such systems, but I think in practice, a significant part is going to be like you are doing the calculus, or people deploying systems are doing the calculus as they do today, in many cases, overtly of like, look, these systems are not very well controlled or understood.

“You think you can understand or control such systems, but I think in practice, a significant part is going to be like you are doing the calculus, or people deploying systems are doing the calculus as they do today, in many cases, overtly of like, look, these systems are not very well controlled or understood.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

07 / belief

I think how you would automate interpretability if you wanted to right now is you take the process humans use it’s like great, we’re going to take that human process, train ML systems to do the pieces that humans do of that process and then just do a lot more of it.

“I think how you would automate interpretability if you wanted to right now is you take the process humans use it’s like great, we’re going to take that human process, train ML systems to do the pieces that humans do of that process and then just do a lot more of it.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

09 / belief

If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn’t have any concise explanation for the distribution. So this first problem, just like it’s a whole different set of activations, I think you’re actually kind of okay, and then the thing that becomes more messy is like but the real world will not only be new samples of different activations, they will also be different in important ways.

“If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn’t have any concise explanation for the distribution. So this first problem, just like it’s a whole different set of activations, I think you’re actually kind of okay, and then the thing that becomes more messy is like but the real world will not only be new samples of different activations, they will also be different in important ways.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

10 / evaluation

Correct me if this is wrong, but it sounds like you basically give it the opportunity to do a coup or make a bioweapon or whatever in testing in a situation where it thinks it’s the real world and you’re like, it didn’t do any of that.

“Correct me if this is wrong, but it sounds like you basically give it the opportunity to do a coup or make a bioweapon or whatever in testing in a situation where it thinks it’s the real world and you’re like, it didn’t do any of that.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

11 / uncertainty

I mean, it depends a little bit what you mean by succeed. But if you, say, get explanations that are great and accurately reflect reality and work for all of these applications that we’re imagining or that we are optimistic about, like kind of the best case success, I don’t know, like 1020 percent something.

“I mean, it depends a little bit what you mean by succeed. But if you, say, get explanations that are great and accurately reflect reality and work for all of these applications that we’re imagining or that we are optimistic about, like kind of the best case success, I don’t know, like 1020 percent something.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

12 / belief

What is a good explanation? And even when people are doing informal interpretability I think if you’re publishing in an ML conference and you want to say this is a good explanation, the way you would verify that would even if not like a formal set of causal intervention experiments.

“What is a good explanation? And even when people are doing informal interpretability I think if you’re publishing in an ML conference and you want to say this is a good explanation, the way you would verify that would even if not like a formal set of causal intervention experiments.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

13 / belief

I think neural nets are kind of unusual in being a domain where we really do want to do systematic, formal reasoning, even though we’re not trying to get a lot of confidence, we’re just trying to understand even roughly what’s going on.

“I think neural nets are kind of unusual in being a domain where we really do want to do systematic, formal reasoning, even though we’re not trying to get a lot of confidence, we’re just trying to understand even roughly what’s going on.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

16 / belief

I think ultimately those have to be international agreements and you might hope they're made more danger by danger, but you might also make them in a very broad way with respect to AI.

“I think ultimately those have to be international agreements and you might hope they're made more danger by danger, but you might also make them in a very broad way with respect to AI.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

17 / uncertainty

You could ask what’s the shortest program, which if you run it on a million h 100s connected in a nice network with a hospitable environment will eventually go to the stars. But that seems like it’s probably on the order of tens of thousands of bytes or I don’t know if I had to guess the median, I’d guess 10,000 bytes.

“You could ask what’s the shortest program, which if you run it on a million h 100s connected in a nice network with a hospitable environment will eventually go to the stars. But that seems like it’s probably on the order of tens of thousands of bytes or I don’t know if I had to guess the median, I’d guess 10,000 bytes.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

19 / belief

I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment.

“I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

20 / belief

Not all of it, but a significant part. It’s just like the difficulty of formalizing the proof is like the hard part and actually getting all of that to go through and we’re not going to help even the tiniest bit with that, I think.

“Not all of it, but a significant part. It’s just like the difficulty of formalizing the proof is like the hard part and actually getting all of that to go through and we’re not going to help even the tiniest bit with that, I think.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

21 / belief

Neutralize the threat is a similar issue where it's just like, I think you would kill the humans if you didn't care at all about them. So maybe your question you're asking is, like, why would you care at all about but I think you don't have to care much to not kill the humans.

“Neutralize the threat is a similar issue where it's just like, I think you would kill the humans if you didn't care at all about them. So maybe your question you're asking is, like, why would you care at all about but I think you don't have to care much to not kill the humans.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

22 / belief

I think the most useful profile is probably a combination of intellectually interested in this particular project and motivated enough by alignment to work on this project, even if it’s really hard.

“I think the most useful profile is probably a combination of intellectually interested in this particular project and motivated enough by alignment to work on this project, even if it’s really hard.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

23 / belief

I think as you move beyond that, it really depends how you’re deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weights sitting there, I think you mostly have addressed the risk from the weights just sitting there.

“I think as you move beyond that, it really depends how you’re deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weights sitting there, I think you mostly have addressed the risk from the weights just sitting there.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

24 / uncertainty

Yeah, I mean, like, for example, if you imagine what the best thing is, it would almost certainly involve just like simulating every possible universe. It might be in modular moral constraints, which I don’t know if you want to include like so that would be very slow.

“Yeah, I mean, like, for example, if you imagine what the best thing is, it would almost certainly involve just like simulating every possible universe. It might be in modular moral constraints, which I don’t know if you want to include like so that would be very slow.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

25 / belief

I'm just saying, okay, I personally, it's hard for me to imagine in 100 years that these things are still our slaves. And if they are, I think that's not the best world.

“I'm just saying, okay, I personally, it's hard for me to imagine in 100 years that these things are still our slaves. And if they are, I think that's not the best world.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

26 / commitment

Maybe a higher level thing that goes into both of these. And then I will talk about how you instantiate an a causal trade is just like it matters a lot to the humans not to get murdered.

“Maybe a higher level thing that goes into both of these. And then I will talk about how you instantiate an a causal trade is just like it matters a lot to the humans not to get murdered.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

27 / uncertainty

Like if you set aside manufacturing costs and just look at operating costs or performance trade offs, like, I don't know, more like 3 orders of magnitude or something like that, or some things that.

“Like if you set aside manufacturing costs and just look at operating costs or performance trade offs, like, I don't know, more like 3 orders of magnitude or something like that, or some things that.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

28 / belief

One is like, I don't believe the chimp thing is going to be as abrupt. That is, I think if you scaled up from chimps to humans, you actually see quite large economic value from the fully domesticated chimp already.

“One is like, I don't believe the chimp thing is going to be as abrupt. That is, I think if you scaled up from chimps to humans, you actually see quite large economic value from the fully domesticated chimp already.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

29 / belief

I think the mental reference would be like, I don’t really like proofs because I think there’s such a huge gap between what you can prove and how you would analyze a neural net.

“I think the mental reference would be like, I don’t really like proofs because I think there’s such a huge gap between what you can prove and how you would analyze a neural net.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

30 / belief

We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it’s being trained to thinking it should defect.

“We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it’s being trained to thinking it should defect.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

31 / belief

I think a qualitative consideration that could significantly slow things down is just like right now you get to observe this really rich supervision from basically next word prediction, or in practice, maybe you're looking at a couple of sentences prediction.

“I think a qualitative consideration that could significantly slow things down is just like right now you get to observe this really rich supervision from basically next word prediction, or in practice, maybe you're looking at a couple of sentences prediction.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

32 / uncertainty

Might be like Rlhf where we don’t know if it generalizes, but so far it makes your chat GPT thing better and you can also use it to make sure that chat GPT doesn’t tell you how to make a bioweapon.

“Might be like Rlhf where we don’t know if it generalizes, but so far it makes your chat GPT thing better and you can also use it to make sure that chat GPT doesn’t tell you how to make a bioweapon.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

34 / uncertainty

I don’t know, it’s hundreds of symbols or something that go into the entire foundations and the entire rules of reasoning for like there’s a sort of built on top of first order logic but the rules of reasoning for first order logic are just like another hundreds of symbols or 100 lines of code or whatever.

“I don’t know, it’s hundreds of symbols or something that go into the entire foundations and the entire rules of reasoning for like there’s a sort of built on top of first order logic but the rules of reasoning for first order logic are just like another hundreds of symbols or 100 lines of code or whatever.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

35 / belief

I think speeding up locally is a little bit less rough. And then, yeah, I think that the effect, like the overall effect size from doing alignment work on reducing takeover risk versus speeding up AI is pretty good.

“I think speeding up locally is a little bit less rough. And then, yeah, I think that the effect, like the overall effect size from doing alignment work on reducing takeover risk versus speeding up AI is pretty good.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

36 / commitment

I’ll set out to be interested in fundraising and someone will be like, offer a grant, and then I will get to delay for another six months or fundraising or nine months, or you can you can delay the time at which Paul needs to think for some time about fundraising.

“I’ll set out to be interested in fundraising and someone will be like, offer a grant, and then I will get to delay for another six months or fundraising or nine months, or you can you can delay the time at which Paul needs to think for some time about fundraising.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

37 / belief

Like, large language models look really good on paper, and RLH looks really good on paper. And these things, I think, just work out in a way that’s yeah, I think people maybe overestimate or, like, maybe it’s kind of a trope, but people talk about, like, it’s easy to underestimate how much gap there is to practice, like, how many things will come up that don’t come up in theory.

“Like, large language models look really good on paper, and RLH looks really good on paper. And these things, I think, just work out in a way that’s yeah, I think people maybe overestimate or, like, maybe it’s kind of a trope, but people talk about, like, it’s easy to underestimate how much gap there is to practice, like, how many things will come up that don’t come up in theory.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

38 / belief

Like, I think people there’s a reasonable chance if you don’t have a splash about chat GBT, you have a splash about GBT four, and if you fail to have a splash about GBT four, there’s a reasonable chance of a splash about GBT 4.

“Like, I think people there’s a reasonable chance if you don’t have a splash about chat GBT, you have a splash about GBT four, and if you fail to have a splash about GBT four, there’s a reasonable chance of a splash about GBT 4.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

41 / belief

Advisors helping them run persuasion campaigns or whatever. But anyway, I think for the most part the default remedy is think about particular harms, have legal protections either in the use of physical technologies that are relevant or in access to AI advice or whatever else to protect against those harms.

“Advisors helping them run persuasion campaigns or whatever. But anyway, I think for the most part the default remedy is think about particular harms, have legal protections either in the use of physical technologies that are relevant or in access to AI advice or whatever else to protect against those harms.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

43 / belief

You'd prefer be able to determine your own destiny, control your competing hardware, et cetera, which I think probably emerge a little bit later than systems that try and get reward and so will generalize in scary, unpredictable ways to new situations.

“You'd prefer be able to determine your own destiny, control your competing hardware, et cetera, which I think probably emerge a little bit later than systems that try and get reward and so will generalize in scary, unpredictable ways to new situations.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

44 / evaluation

It's not the case that humans could defend them from an invasion on their own. So that is if you had invading army and you had your own robot army, you can't just be like, we're going to turn off the robots now because things are going wrong if you're in the middle of a war.

“It's not the case that humans could defend them from an invasion on their own. So that is if you had invading army and you had your own robot army, you can't just be like, we're going to turn off the robots now because things are going wrong if you're in the middle of a war.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

45 / prediction

I think you probably get like, by 2040, like, I don't know, 3 orders of magnitude of effective training compute improvement or like, a good chunk of effective training compute improvement, 4 orders of magnitude.

“I think you probably get like, by 2040, like, I don't know, 3 orders of magnitude of effective training compute improvement or like, a good chunk of effective training compute improvement, 4 orders of magnitude.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

46 / evaluation

You say like, look, a matter of policy sort of access to industry is somewhat restricted or somewhat regulated, even though, again, right now it can be mostly regulated just because most people aren't rich enough that they could even go off and just build 1000 tanks.

“You say like, look, a matter of policy sort of access to industry is somewhat restricted or somewhat regulated, even though, again, right now it can be mostly regulated just because most people aren't rich enough that they could even go off and just build 1000 tanks.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

47 / evaluation

Like, maybe most saliently here, if you want to use some biological weapons or some crazy shit that might just kill humans, I think you might kill humans just from totally destroying the ecosystems they're dependent on.

“Like, maybe most saliently here, if you want to use some biological weapons or some crazy shit that might just kill humans, I think you might kill humans just from totally destroying the ecosystems they're dependent on.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

49 / preference

I tend to be a little bit more skeptical about those arguments and tend to think, like, yeah, something can be bullshit because it’s not addressing a real problem that’s I think the easiest way this is a problem someone’s interested in that’s just not actually an important problem, and there’s no story about why it’s going to become an important problem.

“I tend to be a little bit more skeptical about those arguments and tend to think, like, yeah, something can be bullshit because it’s not addressing a real problem that’s I think the easiest way this is a problem someone’s interested in that’s just not actually an important problem, and there’s no story about why it’s going to become an important problem.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

50 / prediction

The good story is you develop methods that address a bunch of existing problems because they just are more principled ways to train AI systems that work better, people adopt them and then we are no longer worried about eg reward hacking or deceptive alignment.

“The good story is you develop methods that address a bunch of existing problems because they just are more principled ways to train AI systems that work better, people adopt them and then we are no longer worried about eg reward hacking or deceptive alignment.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

51 / evaluation

It’s not in control of factories and robot armies or whatever. So in that case, even in training it will have those activations for being fucked up on because in the back of its mind it’s thinking I will take over once I have the opportunity.

“It’s not in control of factories and robot armies or whatever. So in that case, even in training it will have those activations for being fucked up on because in the back of its mind it’s thinking I will take over once I have the opportunity.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

52 / prediction

First things would go badly without it. But I think if you ask, why don't we turn off the AI, my best guess is because there are a bunch of other AIS running around 2D or lunch.

“First things would go badly without it. But I think if you ask, why don't we turn off the AI, my best guess is because there are a bunch of other AIS running around 2D or lunch.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

53 / prediction

I think once you have systems actually doing jobs, the extrapolation gets easier because you're not moving from a subjective impression of a chat to automating all R and D, you're moving from automating this job to automating that job or whatever.

“I think once you have systems actually doing jobs, the extrapolation gets easier because you're not moving from a subjective impression of a chat to automating all R and D, you're moving from automating this job to automating that job or whatever.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast
Search evidence