High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Lex Fridman: belief

2 Jun 2024 Lex Fridman Podcast #431 – Roman Yampolskiy: Dangers of Superintelligent AI

“I think we’ll just get a lot of examples of deceptions from large language models or AI systems.”

— Lex Fridman

Source trail

Everything needed to verify it.

Speaker
Lex Fridman
Attribution
Verified speaker
Claim type
belief
Recorded
2 Jun 2024
Publisher
Lex Fridman Podcast

Transcript context

…Correct. I don’t think developers know everything about what they are creating. They have lots of great knowledge, we’re making progress on explaining parts of a network. We can understand, “Okay, this note get excited, then this input is presented, this cluster of notes.” But we’re nowhere near close to understanding the full picture, and I think it’s impossible. You need to be able to survey an explanation. The size of those models prevents a single human from absorbing all this information, even if provided by the system. So either we’re getting model as an explanation for what’s happening and that’s not comprehensible to us or we’re getting compressed explanation, [inaudible 00:59:01] compression, where here, “Top 10 reasons you got fired.” It’s something, but it’s not a full picture. You’ve given elsewhere an example of a child and everybody, all humans try to deceive, they try to lie early on in their life. I think we’ll just get a lot of examples of deceptions from large language models or AI systems. They’re going to be kind of shady, or they’ll be pretty good, but we’ll catch them off guard. We’ll start to see the kind of momentum towards developing increasing deception capabilities and that’s when you’re like, “Okay, we need to do some kind of alignment that prevents deception.” But, if you support open source, then you can have open source models that have some level of deception you can start to explore on a large scale, how do we stop it from being deceptive? Then there’s a more explicit, pragmatic kind of problem to solve. How do we stop AI systems from trying to optimize for deception? That’s an example. So there is a paper, I think it came out last week by Dr Park et al, from MIT I think, and they showed that models already showed successful deception in what they do. My concern is not that they lie now, and we need to catch them and tell them, “Don’t lie.” My concern is that once they are capable and deployed, they will later change their mind. Because what unrestricted learning allows you to do. Lots of people grow up maybe in the religious family, they read some new books and they turn in their religion. That’s a treacherous turn in humans. If you learn something new about your colleagues, maybe you’ll change how you react to that.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence