High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Paul Christiano

Published podcast speaker

Claims
46
Episodes
1
Shows
1
Named items
0

Claim ledger

What Paul said.

26 transcript-backed records

02 / belief

You think you can understand or control such systems, but I think in practice, a significant part is going to be like you are doing the calculus, or people deploying systems are doing the calculus as they do today, in many cases, overtly of like, look, these systems are not very well controlled or understood.

“You think you can understand or control such systems, but I think in practice, a significant part is going to be like you are doing the calculus, or people deploying systems are doing the calculus as they do today, in many cases, overtly of like, look, these systems are not very well controlled or understood.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

04 / belief

I think how you would automate interpretability if you wanted to right now is you take the process humans use it’s like great, we’re going to take that human process, train ML systems to do the pieces that humans do of that process and then just do a lot more of it.

“I think how you would automate interpretability if you wanted to right now is you take the process humans use it’s like great, we’re going to take that human process, train ML systems to do the pieces that humans do of that process and then just do a lot more of it.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

06 / belief

If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn’t have any concise explanation for the distribution. So this first problem, just like it’s a whole different set of activations, I think you’re actually kind of okay, and then the thing that becomes more messy is like but the real world will not only be new samples of different activations, they will also be different in important ways.

“If every data point from the distribution was a whole new thing happening for different reasons, you actually couldn’t have any concise explanation for the distribution. So this first problem, just like it’s a whole different set of activations, I think you’re actually kind of okay, and then the thing that becomes more messy is like but the real world will not only be new samples of different activations, they will also be different in important ways.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

07 / belief

What is a good explanation? And even when people are doing informal interpretability I think if you’re publishing in an ML conference and you want to say this is a good explanation, the way you would verify that would even if not like a formal set of causal intervention experiments.

“What is a good explanation? And even when people are doing informal interpretability I think if you’re publishing in an ML conference and you want to say this is a good explanation, the way you would verify that would even if not like a formal set of causal intervention experiments.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

08 / belief

I think neural nets are kind of unusual in being a domain where we really do want to do systematic, formal reasoning, even though we’re not trying to get a lot of confidence, we’re just trying to understand even roughly what’s going on.

“I think neural nets are kind of unusual in being a domain where we really do want to do systematic, formal reasoning, even though we’re not trying to get a lot of confidence, we’re just trying to understand even roughly what’s going on.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

10 / belief

I think ultimately those have to be international agreements and you might hope they're made more danger by danger, but you might also make them in a very broad way with respect to AI.

“I think ultimately those have to be international agreements and you might hope they're made more danger by danger, but you might also make them in a very broad way with respect to AI.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

12 / belief

I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment.

“I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment, then there would just be no possible way that the calculus comes out negative for alignment.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

13 / belief

Not all of it, but a significant part. It’s just like the difficulty of formalizing the proof is like the hard part and actually getting all of that to go through and we’re not going to help even the tiniest bit with that, I think.

“Not all of it, but a significant part. It’s just like the difficulty of formalizing the proof is like the hard part and actually getting all of that to go through and we’re not going to help even the tiniest bit with that, I think.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

14 / belief

Neutralize the threat is a similar issue where it's just like, I think you would kill the humans if you didn't care at all about them. So maybe your question you're asking is, like, why would you care at all about but I think you don't have to care much to not kill the humans.

“Neutralize the threat is a similar issue where it's just like, I think you would kill the humans if you didn't care at all about them. So maybe your question you're asking is, like, why would you care at all about but I think you don't have to care much to not kill the humans.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

15 / belief

I think the most useful profile is probably a combination of intellectually interested in this particular project and motivated enough by alignment to work on this project, even if it’s really hard.

“I think the most useful profile is probably a combination of intellectually interested in this particular project and motivated enough by alignment to work on this project, even if it’s really hard.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

16 / belief

I think as you move beyond that, it really depends how you’re deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weights sitting there, I think you mostly have addressed the risk from the weights just sitting there.

“I think as you move beyond that, it really depends how you’re deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weights sitting there, I think you mostly have addressed the risk from the weights just sitting there.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

17 / belief

One is like, I don't believe the chimp thing is going to be as abrupt. That is, I think if you scaled up from chimps to humans, you actually see quite large economic value from the fully domesticated chimp already.

“One is like, I don't believe the chimp thing is going to be as abrupt. That is, I think if you scaled up from chimps to humans, you actually see quite large economic value from the fully domesticated chimp already.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

18 / belief

I think the mental reference would be like, I don’t really like proofs because I think there’s such a huge gap between what you can prove and how you would analyze a neural net.

“I think the mental reference would be like, I don’t really like proofs because I think there’s such a huge gap between what you can prove and how you would analyze a neural net.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

19 / belief

We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it’s being trained to thinking it should defect.

“We believe that adversarial training fixes this, or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it’s being trained to thinking it should defect.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

20 / belief

I think a qualitative consideration that could significantly slow things down is just like right now you get to observe this really rich supervision from basically next word prediction, or in practice, maybe you're looking at a couple of sentences prediction.

“I think a qualitative consideration that could significantly slow things down is just like right now you get to observe this really rich supervision from basically next word prediction, or in practice, maybe you're looking at a couple of sentences prediction.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

21 / belief

I think speeding up locally is a little bit less rough. And then, yeah, I think that the effect, like the overall effect size from doing alignment work on reducing takeover risk versus speeding up AI is pretty good.

“I think speeding up locally is a little bit less rough. And then, yeah, I think that the effect, like the overall effect size from doing alignment work on reducing takeover risk versus speeding up AI is pretty good.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

22 / belief

Like, large language models look really good on paper, and RLH looks really good on paper. And these things, I think, just work out in a way that’s yeah, I think people maybe overestimate or, like, maybe it’s kind of a trope, but people talk about, like, it’s easy to underestimate how much gap there is to practice, like, how many things will come up that don’t come up in theory.

“Like, large language models look really good on paper, and RLH looks really good on paper. And these things, I think, just work out in a way that’s yeah, I think people maybe overestimate or, like, maybe it’s kind of a trope, but people talk about, like, it’s easy to underestimate how much gap there is to practice, like, how many things will come up that don’t come up in theory.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

23 / belief

Like, I think people there’s a reasonable chance if you don’t have a splash about chat GBT, you have a splash about GBT four, and if you fail to have a splash about GBT four, there’s a reasonable chance of a splash about GBT 4.

“Like, I think people there’s a reasonable chance if you don’t have a splash about chat GBT, you have a splash about GBT four, and if you fail to have a splash about GBT four, there’s a reasonable chance of a splash about GBT 4.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

24 / belief

Advisors helping them run persuasion campaigns or whatever. But anyway, I think for the most part the default remedy is think about particular harms, have legal protections either in the use of physical technologies that are relevant or in access to AI advice or whatever else to protect against those harms.

“Advisors helping them run persuasion campaigns or whatever. But anyway, I think for the most part the default remedy is think about particular harms, have legal protections either in the use of physical technologies that are relevant or in access to AI advice or whatever else to protect against those harms.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast

26 / belief

You'd prefer be able to determine your own destiny, control your competing hardware, et cetera, which I think probably emerge a little bit later than systems that try and get reward and so will generalize in scary, unpredictable ways to new situations.

“You'd prefer be able to determine your own destiny, control your competing hardware, et cetera, which I think probably emerge a little bit later than systems that try and get reward and so will generalize in scary, unpredictable ways to new situations.”
Speaker
Paul Christiano
Publisher
Dwarkesh Podcast
Search evidence