Evidence receipt / evaluation
Published · transcript-backedJohn Schulman: evaluation
15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
“We found that a tiny amount of data did the trick, even when you mixed it together with everything else.”
Source trail
Everything needed to verify it.
- Speaker
- John Schulman
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…I guess that's been the big lesson of language models, right? That they can act within the similar affordances that humans do. I want to go back to the point you made earlier about how this process could be more sample efficient because it could generalize from its pre-training experiences of how to get unstuck in different scenarios. What is the strongest evidence you’ve seen of this generalization and transfer? The big question for the future abilities models seems to be about how much generalization is happening. Is there something that feels really compelling to you? Have you seen a model learn something that you wouldn't expect it to learn from generalization? There have definitely been some interesting instances of generalization in post-training. One well-known phenomenon is that if you do all your fine-tuning with English data, the model will automatically behave well in other languages. So if you train the assistant on English data, it'll also do something reasonable in Spanish. Sometimes you might get the wrong behavior in terms of whether it replies in English or replies in Spanish. Usually you get the right behavior there, meaning you get it to respond in Spanish to Spanish queries. That's one interesting instance of generalization where the model just latches onto the right, helpful persona and then automatically does the right thing in different languages. We've seen some version of this with multimodal data where if you do text-only fine-tuning, you also get reasonable behavior with images. Early on in ChatGPT we were trying to fix some issues with the model understanding its own limitations. Early versions of the model would think that it could send you an email or call you an Uber or something. The model would try to play the assistant and it would say “oh yeah, of course I sent that email.” Obviously it didn't. So we started collecting some data to fix those problems. We found that a tiny amount of data did the trick, even when you mixed it together with everything else. I don't remember exactly how many examples but something like 30 examples. We had a pretty small number of examples showing this general behavior, explaining that the model doesn’t have this capability. That generalized pretty well to all sorts of capabilities we didn't train for. I still want to go back to this because I'm not sure I understood. Let’s say you have this model that is trained to be coherent for longer periods of time. Setting aside these other bottlenecks which there may or may not be, by next could you have models that are potentially like human-level? I’m thinking of a model that you’re interacting with like a colleague and it's as good as interacting with a human colleague. You can tell them to go do stuff and they go get it done. What seems wrong with that picture of the capabilities you think might be possible?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.