Evidence receipt / evaluation
Published · transcript-backedRoman Yampolskiy: evaluation
2 Jun 2024 Lex Fridman Podcast #431 – Roman Yampolskiy: Dangers of Superintelligent AI
“The problem is that it’s very hard to separate capabilities work from safety work.”
Source trail
Everything needed to verify it.
- Speaker
- Roman Yampolskiy
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 2 Jun 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…Right. Is there any actual explicit capabilities that you can put on paper, that we as a human civilization could put on paper? Is it possible to make it explicit like that versus kind of a vague notion of just like you said, it’s very vague. We want AI systems to do good and want them to be safe. Those are very vague notions. Is there more formal notions? So, when I think about this problem, I think about having a toolbox I would need. Capabilities such as explaining everything about that system’s design and workings, predicting not just terminal goal, but all the intermediate steps of a system. Control in terms of either direct control, some sort of a hybrid option, ideal advisor. It doesn’t matter which one you pick, but you have to be able to achieve it. In a book we talk about others, verification is another very important tool. Communication without ambiguity, human language is ambiguous. That’s another source of danger. So, basically there is a paper we published in ACM surveys, which looks at about 50 different impossibility results, which may or may not be relevant to this problem, but we don’t have enough human resources to investigate all of them for relevance to AI safety. The ones I mentioned to you, I definitely think would be handy, and that’s what we see AI safety researchers working on. Explainability is a huge one. The problem is that it’s very hard to separate capabilities work from safety work. If you make good progress in explainability, now the system itself can engage in self-improvement much easier, increasing capability greatly. So, it’s not obvious that there is any research which is pure safety work without disproportionate increasing capability and danger. Explainability is really interesting. Why is that connected to you to capability? If it’s able to explain itself well, why does that naturally mean that it’s more capable?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.