High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

John Schulman: belief

15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

“Obviously, the user might ask the model to do something that we think is actively harmful to other people.”

— John Schulman

Source trail

Everything needed to verify it.

Speaker
John Schulman
Attribution
Verified speaker
Claim type
belief
Recorded
15 May 2024
Publisher
Dwarkesh Podcast

Transcript context

…So this is a problem we have to deal with today with RLHF. You have to aggregate preferences across a lot of different humans. It'll be maybe more marked with future, more powerful systems. But when you say we want these eventual AI systems that are going to fully replace humans as part of these firms to be aligned, what does that mean? Will it mean that they basically do what the user wants them to do? Does it mean that they have to result in some sort of global outcome that we're happy with as the stakeholders in OpenAI? Concretely, what would that mean? If the models are being used for these higher stakes use cases, then we would have to think about RLHF in a much different way than we are right now.We're not quite ready for that or the current methods might not be completely sufficient. We would need to make compromises between the needs of the different stakeholders involved. We have this document that we're releasing called the Model Spec. It's about how we want our models to behave in the API and in ChatGPT. We try to talk about this issue where there are different stakeholders involved and sometimes there are conflicts between what they might want. In our case, we were thinking of the stakeholders as the end user (someone sitting in front of ChatGPT or some other app), the developer (someone using the API who might be serving other end users with their app), the platform (OpenAI, we don't want the models to expose us to legal risk), and the rest of humanity (including people not part of the users or customers). Obviously, the user might ask the model to do something that we think is actively harmful to other people. We might have to refuse that. By the way, this isn't the order of priority necessarily. These are just the four or so classes of stakeholder. Actually, you could maybe also say in the future, the model itself. We're not there yet. Anyway, we have these different stakeholders. Sometimes they have conflicting demands. We have to make some call on how to resolve those conflicts.It's not always obvious how to do that. We had to think through the trade-offs and basically the rough heuristic is that we mostly want the models to follow your instructions and be helpful to the user and the developer. But when this impinges on other people's happiness or way of life, this becomes a problem and we have to block certain kinds of usage. We mostly want the models to just be an extension of people's will and do what they say. We don't want to be too paternalistic. We want to be neutral and not impose our opinions on people. We mostly want to let people do what they want with the models. I got a chance to read the Spec beforehand. This is a question of how well that transfers over to how the model itself behaves. I was impressed with how sensible the trade-offs were. I believe the actual edge cases were explicitly stated rather than the kinds of things where are obvious. In this case, you really are going after the edge cases.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence