Evidence receipt / evaluation
Published · transcript-backedEliezer Yudkowsky: evaluation
6 Apr 2023 Dwarkesh Podcast Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality
“Getting stuff to be more likely does not help you if the baseline is nearly zero.”
Source trail
Everything needed to verify it.
- Speaker
- Eliezer Yudkowsky
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 6 Apr 2023
- Publisher
- Dwarkesh Podcast
Transcript context
…Would you at least say that we are living in a better situation than one in which we have some sort of black box where you have a machiavellian fittest survive simulation that produces AI? This situation is at least more likely to produce alignment than one in which something that is completely untouched by human psychology would produce? More likely? Yes. Maybe you’re an order of magnitude likelier. 0% instead of 0%. Getting stuff to be more likely does not help you if the baseline is nearly zero. The whole training set up there is producing an actress, a predictor. It’s not actually being put into the kind of ancestral situation that evolved humans, nor the kind of modern situation that raises humans. Though to be clear, raising it like a human wouldn’t help, But you’re giving it a very alien problem that is not what humans solve and it is solving that problem not in the way a human would. Okay, so how about this. I can see that I certainly don’t know for sure what is going on in these systems. In fact, obviously nobody does. But that also goes through you. Could it not just be that reinforcement learning works and all these other things we’re trying somehow work and actually just being an actor produces some sort of benign outcome where there isn’t that level of simulation and conniving?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.