High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Paul Christiano: belief

31 Oct 2023 Dwarkesh Podcast Paul Christiano — Preventing an AI takeover

“Like, large language models look really good on paper, and RLH looks really good on paper. And these things, I think, just work out in a way that’s yeah, I think people maybe overestimate or, like, maybe it’s kind of a trope, but people talk about, like, it’s easy to underestimate how much gap there is to practice, like, how many things will come up that don’t come up in theory.”

— Paul Christiano

Source trail

Everything needed to verify it.

Speaker
Paul Christiano
Attribution
Verified speaker
Claim type
belief
Recorded
31 Oct 2023
Publisher
Dwarkesh Podcast

Transcript context

…Well, one question I think would be interesting to ask, you know, I think people can talk vaguely about the value of theoretical research and how it contributes to real world applications and you can look at historical examples or something, but you are somebody who actually has done this in a big way. Like Rlhf is something you developed and then it actually has got into an application that has been used by millions of people. Tell me about just that pipeline. How can you reliably identify theoretical problems that will matter for real world applications? Because it’s one thing to read about touring or something and the Halting problem, but here you’d have the real thing. Yeah, I mean, it is definitely exciting to have worked on a thing that has a real world impact. The main caveat I’d provide is, like, Rlhf is very simple compared to many things. And so the motivation for working on that problem was, like, look, this is how it probably should work, or this is a step in some progression. It’s unclear if it’s, like, the final step or something, but it’s a very natural thing to do that people probably should be and probably will be doing. I’m saying, if you want to talk about crazy stuff, it’s good to help make those steps happen faster, and it’s good to learn about. There’s lots of issues that occur in practice, even for things that seem very simple on paper, but mostly, like, the story of it’s just like, yeah, I think my sense of the world is things that look like good ideas on paper, just, like, often are harder than they look. But the world isn’t that far from what makes sense on paper. Like, large language models look really good on paper, and RLH looks really good on paper. And these things, I think, just work out in a way that’s yeah, I think people maybe overestimate or, like, maybe it’s kind of a trope, but people talk about, like, it’s easy to underestimate how much gap there is to practice, like, how many things will come up that don’t come up in theory. But it’s also easy to overestimate how inscrutable the world is. Like, the things that happen mostly are things that do just kind of make sense. Yeah, I feel like most ML implementation does just come down to a bunch of detail, though, of, like, build a very simple version of the system, understand what goes wrong, fix the things that go wrong, scale it up, understand what goes wrong. And I’m glad I have some experience doing that, but I think that does cause me to be better informed about what makes sense in ML and what can actually work. But I don’t think it caused me to have a whole lot of deep expertise or deep wisdom about how to close the gap. Yeah, but is there some tip on identifying things like Rlhf which actually do matter, versus making sure you don’t get stuck in some theoretical problem that doesn’t matter? Or is it just coincidence? Or I mean, is there something you can do in advance to make sure that the thing is useful?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence