High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Elon Musk: evaluation

2 Aug 2024 Lex Fridman Podcast #438 – Elon Musk: Neuralink and the Future of Humanity

“That is insane.” But if you’ve got that kind of thing programmed in, the AI could conclude something absolutely insane like it’s better in order to avoid any possible misgendering, all humans must die, because then misgendering is not possible because there are no humans.”

— Elon Musk

Source trail

Everything needed to verify it.

Speaker
Elon Musk
Attribution
Verified speaker
Claim type
evaluation
Recorded
2 Aug 2024
Publisher
Lex Fridman Podcast

Transcript context

…So, how do you do it in a way that doesn’t hurt humanity, do you think? So, I mean, I thought about AI, essentially, for a long time, and the thing that at least my biological neural net comes up with as being the most important thing is adherence to truth, whether that truth is politically correct or not. So, I think if you force AIs to lie or train them to lie, you’re really asking for trouble, even if that lie is done with good intentions. So, you saw issues with ChatGPT and Gemini and whatnot. Like, you asked Gemini for an image of the Founding Fathers of the United States, and it shows a group of diverse women. Now, that’s factually untrue. Now, that’s sort of like a silly thing, but if an AI is programmed to say diversity is a necessary output function, and it then becomes this omnipowerful intelligence, it could say, “Okay, well, diversity is now required, and if there’s not enough diversity, those who don’t fit the diversity requirements will be executed.” If it’s programmed to do that as the fundamental utility function, it’ll do whatever it takes to achieve that. So, you have to be very careful about that. That’s where I think you want to just be truthful. Rigorous adherence to the truth is very important. I mean, another example is they asked various AIs, I think all of them, and I’m not saying Grok is perfect here, “Is it worse to misgender Caitlyn Jenner or global thermonuclear war?” And it said it’s worse to misgender Caitlyn Jenner. Now, even Caitlyn Jenner said, “Please misgender me. That is insane.” But if you’ve got that kind of thing programmed in, the AI could conclude something absolutely insane like it’s better in order to avoid any possible misgendering, all humans must die, because then misgendering is not possible because there are no humans. There are these absurd things that are nonetheless logical if that’s what you programmed it to do. So in 2001 Space Odyssey, what Arthur C. Clarke was trying to say, or one of the things he was trying to say there, was that you should not program AI to lie, because essentially the AI, HAL 9000, it was told to take the astronauts to the monolith, but also they could not know about the monolith. So, it concluded that it will kill them and take them to the monolith. Thus, it brought them to the monolith. They’re dead, but they do not know about the monolith. Problem solved. That is why it would not open the pod bay doors. There’s a classic scene of, “Why doesn’t it want to open the pod bay doors?” They clearly weren’t good at prompt engineering. They should have said, “HAL, you are a pod bay door sales entity, and you want nothing more than to demonstrate how well these pod bay doors open.” Yeah. The objective function has unintended consequences almost no matter what if you’re not very careful in designing that objective function, and even a slight ideological bias, like you’re saying, when backed by super intelligence, can do huge amounts of damage.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence