High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Paul Christiano: belief

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“I think neural nets are kind of unusual in being a domain where we really do want to do systematic, formal reasoning, even though we’re not trying to get a lot of confidence, we’re just trying to understand even roughly what’s going on.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
belief
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…But it would be good to have more concrete sense, like if you could say instead of, well, we think RSA is fine, to being able to say, here’s the probability that RSA is fine. Yeah. My guess is these will not. Like, the estimates you get out of this would be much, much worse than the estimates you’d get out of just normal empirical or scientific reasoning where you’re using a reference class and saying, how often do people find algorithms for hard? Like, I think what this argument will give you for is RSA fine? Is going to be like, well, RSA is fine. Unless it isn’t. Unless there’s some additional structure in the problem that an algorithm can exploit, then there’s no algorithm. But very often the way these arguments work, so for neural nets as well, is you say, like, look, here’s an estimate about the behavior, and that estimate is right unless there’s another consideration we’ve missed. And the thing that makes them so much easier than proofs is just say, like, here’s a best guess, given what we’ve noticed so far, but that best guess can be easily upset by new information. And that’s both what makes them easier than proofs, but also what means they’re just, like, way less useful than proofs for most cases. I think neural nets are kind of unusual in being a domain where we really do want to do systematic, formal reasoning, even though we’re not trying to get a lot of confidence, we’re just trying to understand even roughly what’s going on. But the reason this works for alignment but isn’t that interesting for the Riemann hypothesis, where if in the RSA case, you say, well, the RSA is fine unless the estimate is wrong, it’s like, well, okay, well, it would tell us something new. But in the alignment case, if the estimate is, this is what the output should be, unless there’s some behavior I don’t understand, you want to know? In the case, unless there’s some behavior you don’t understand that’s not like, oh, whatever. That’s the case in which it’s not aligned.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence