High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Roman Yampolskiy: evaluation

2 Jun 2024 Lex Fridman Podcast #431 – Roman Yampolskiy: Dangers of Superintelligent AI

“The problem is that it’s very hard to separate capabilities work from safety work.”

— Roman Yampolskiy

Source trail

Everything needed to verify it.

Speaker
Roman Yampolskiy
Attribution
Verified speaker
Claim type
evaluation
Recorded
2 Jun 2024
Publisher
Lex Fridman Podcast

Transcript context

…Right. Is there any actual explicit capabilities that you can put on paper, that we as a human civilization could put on paper? Is it possible to make it explicit like that versus kind of a vague notion of just like you said, it’s very vague. We want AI systems to do good and want them to be safe. Those are very vague notions. Is there more formal notions? So, when I think about this problem, I think about having a toolbox I would need. Capabilities such as explaining everything about that system’s design and workings, predicting not just terminal goal, but all the intermediate steps of a system. Control in terms of either direct control, some sort of a hybrid option, ideal advisor. It doesn’t matter which one you pick, but you have to be able to achieve it. In a book we talk about others, verification is another very important tool. Communication without ambiguity, human language is ambiguous. That’s another source of danger. So, basically there is a paper we published in ACM surveys, which looks at about 50 different impossibility results, which may or may not be relevant to this problem, but we don’t have enough human resources to investigate all of them for relevance to AI safety. The ones I mentioned to you, I definitely think would be handy, and that’s what we see AI safety researchers working on. Explainability is a huge one. The problem is that it’s very hard to separate capabilities work from safety work. If you make good progress in explainability, now the system itself can engage in self-improvement much easier, increasing capability greatly. So, it’s not obvious that there is any research which is pure safety work without disproportionate increasing capability and danger. Explainability is really interesting. Why is that connected to you to capability? If it’s able to explain itself well, why does that naturally mean that it’s more capable?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence