High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality

“I think the credit you would get for that, rightly, is as a good Agnostic forecaster, as somebody who is calm and measured. But it seems like to be able to make really strong claims about the future, about something that is so out of prior distributions as like the death of humanity, you don’t only have to show yourself as a good Agnostic forecaster, you have to show that your ability to forecast because of a particular theory is much greater.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
6 Apr 2023
Publisher
Dwarkesh Podcast

Transcript context

…I feel like a whole bunch of my successful predictions in this have come from other people being like — “Oh, yes. I have this theory which predicts that stuff is 30 years off.” And I’m like — “You don’t know that.” And then stuff happens not 30 years off. And I’m like — “Ha ha. Successful prediction.” And that’s basically what I told you, right? I was like — well, you could have the loss function continuing on a smooth line and new abilities appear, and you could have them suddenly appear in a cluster. Because why not? Because nature just tells you that’s up. And suddenly you can have this one key ability, that’s equivalent to language for humans, and there’s a sudden jump in output capabilities. You could have a new innovation, like the transformer, and maybe the losses actually drop precipitously and a whole bunch of new abilities appear at once. This is all just me. This is me saying — I don’t know. But so many people around are saying things that implicitly claim to know more than that, that it can actually start to sound like a startling prediction. This is one of my big secret tricks, actually. People are like — The AI could be good or evil. So it’s like 50-50, right? And I’m actually like — No, we can be ignorant about a wider space than this in which good is actually like a fairly narrow range. So many of the predictions like that are really anti-predictions. It’s somebody thinking along a relatively narrow line and you point out everything outside of that and it sounds like a startling prediction. Of course, the trouble being, when you look back afterwards, people are like — “Well, those people saying the narrow thing were just silly. Ha ha.” and they don’t give you as much credit. I think the credit you would get for that, rightly, is as a good Agnostic forecaster, as somebody who is calm and measured. But it seems like to be able to make really strong claims about the future, about something that is so out of prior distributions as like the death of humanity, you don’t only have to show yourself as a good Agnostic forecaster, you have to show that your ability to forecast because of a particular theory is much greater. Do you see what I mean? It’s all about the ignorance prior. It’s all about knowing the space in which to be maximum entropy. What will the future be? I don’t know. It could be paperclips, it could be staples. It could be no kind of office supplies at all and tiny little spirals. It could be little tiny things that are like outputting 111, because that’s like the most predictable kind of text to predict. Or representations of ever larger numbers in the fast growing hierarchy because that’s how they interpret the reward counter. I’m actually getting into specifics here, which is the opposite of the point I originally meant to make, which is if somebody claims to be very unsure, I might say — “Okay, so then you expect most possible molecular configurations of the solar system to be equally probable.” Well, humans mostly aren’t in those. So being very unsure about the future looks like predicting with probability nearly one that the humans are all gone, which it’s not actually that bad, but it illustrates the point of people going like — “But how are you sure?” Kind of missing the real discourse and skill, which is like — “Oh, yes, we’re all very unsure. Lots of entropy in our probability distributions. But what is the space under which you are unsure?”…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence