Evidence receipt / belief
Published · transcript-backedRoman Yampolskiy: belief
2 Jun 2024 Lex Fridman Podcast #431 – Roman Yampolskiy: Dangers of Superintelligent AI
“If we had capability of ahead of the run, before the training run to register exactly what capabilities that next model will have at the end of the training run, and we accurately guessed all of them, I would say you’re right, “We can definitely go ahead with this run.”
Source trail
Everything needed to verify it.
- Speaker
- Roman Yampolskiy
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 2 Jun 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…At the training stage, but then there’s a testing stage inside the company, they can start getting intuition about what the system is capable to do. You’re saying that somehow from leap from GPT-4 to GPT-5 can happen, the kind of leap where GPT-4 was controllable and GPT-5 is no longer controllable and we get no insights from using GPT-4 about the fact that GPT-5 will be uncontrollable. That’s the situation you’re concerned about. Where there leap from N, to N plus one will be such that an uncontrollable system is created without any ability for us to anticipate that. If we had capability of ahead of the run, before the training run to register exactly what capabilities that next model will have at the end of the training run, and we accurately guessed all of them, I would say you’re right, “We can definitely go ahead with this run. ” We don’t have the capability. From GPT-4, you can build up intuitions about what GPT-5 will be capable of. It’s just incremental progress. Even if that’s a big leap in capability, it just doesn’t seem like you can take a leap from a system that’s helping you write emails to a system that’s going to destroy human civilization. It seems like it’s always going to be sufficiently incremental such that we can anticipate the possible dangers, and we’re not even talking about existential risk, but just the kind of damage you can do to civilization. It seems like we’ll be able to anticipate the kinds, not the exact, but the kinds of risks it might lead to and then rapidly develop defenses ahead of time and as the risks emerge.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.