High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Paul Christiano: preference

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“I tend to be a little bit more skeptical about those arguments and tend to think, like, yeah, something can be bullshit because it’s not addressing a real problem that’s I think the easiest way this is a problem someone’s interested in that’s just not actually an important problem, and there’s no story about why it’s going to become an important problem.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
preference
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…People have explained their schemes to me and I’m like, honestly, I don’t know if it makes sense or not with you. I’m just like I trust Paul enough that I think there’s probably something here if I try to understand this enough. But how do you detect bullshit? Yeah, so I think it’s depends on the kind of work. So for the kind of stuff we’re doing, my guess is like most people there’s just not really a way you’re going to tell whether it’s bullshit. So I think it’s important that we don’t spend that much money on the people we want to hire are probably going to dig in in depth. I don’t think there’s a way you can tell whether it’s bullshit without either spending a lot of effort or leaning on deference with empirical work. It’s interesting in that you do have some signals of the quality of work. You can be like, does it work in practice? Does the story? I think the stories are just radically simpler and so you probably can evaluate those stories on their face. And then you mostly come down to these questions of like, what are the key difficulties? Yeah, I tend to be optimistic when people dismiss something because this doesn’t deal with a key difficulty or this runs into the following insurable obstacle. I tend to be a little bit more skeptical about those arguments and tend to think, like, yeah, something can be bullshit because it’s not addressing a real problem that’s I think the easiest way this is a problem someone’s interested in that’s just not actually an important problem, and there’s no story about why it’s going to become an important problem. E g, like it’s not a problem now and won’t get worse or it is maybe a problem now, but it’s clearly getting better. That’s like one way and then conditioned on passing that bar, like dealing with something that actually engages with important parts of the argument for concern and then actually making sense empirically. So I think most work is anchored by source of feedback is like actually engaging with real models. So it’s like, does it make sense how to engage with real models? And does the story about how it deals with key difficulties actually make sense? I’m pretty liberal past there. I think it’s really hard to, like, eg. People look at mechanistic, interpretability and be like, well, this obviously can’t succeed. And I’m like, I don’t know. How can you tell? It obviously can’t succeed. I think it’s reasonable to take total investment in the field. How fast is it making progress? How does that pencil I think most things people work on, though, actually pencil pretty fine. They look like they could be reasonable investments. Things are not, like, super out of whack. Okay, great. This is, I think, a good place to close. Paul, thank you so much for your time.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence