Evidence receipt / prediction
Published · transcript-backedDwarkesh Patel: prediction
15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
“Because there doesn't seem to be a model since GPT-4 that seems to be significantly better, there's a hypothesis that we might be hitting some sort of plateau.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 15 May 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…I have a decent amount of experience at this point from the different parts of the stack, from RL algorithms, which I've worked on since grad school, to data collection, annotation processes, and playing with language models. I'd say I've dabbled with these things and the people who do well at this kind of research have some view of the whole stack and have a lot of curiosity about the different parts of it. You want to be both empirical and let experiments update your views, but you also want to think from first principles. Assuming that learning works, what would be the ideal type of data to collect? That type of thing. Because there doesn't seem to be a model since GPT-4 that seems to be significantly better, there's a hypothesis that we might be hitting some sort of plateau. These models aren't actually generalizing that well, and you're going to hit a data wall beyond which the abilities unlocked by memorizing a vast corpus of pre-training data won't help you get something much smarter than GPT-4. Do you think that hypothesis is wrong? We've talked about some examples of generalization, like Spanish to English. One example I think of is the transfer from code to reasoning in language. If you train on a bunch of code, it gets better at reasoning in language? Is that actually the case? Do you see positive transfer between different modalities? If you train on a bunch of videos and images, it'll get smarter from synthetic data? Or does it seem like the abilities unlocked are extremely local to the exact kind of labels and data you put into the training corpus? I'll try to respond to all that. First, are we about to hit the data wall? I wouldn't draw too much from the time since GPT-4 was released because it does take a while to train these models and do all the prep to train a new generation of models. I wouldn't draw too much from that fact. There are definitely some challenges from the limited amount of data, but I wouldn't expect us to immediately hit the data wall. However, I would expect the nature of pre-training to somewhat change over time as we get closer to it. In terms of generalization from different types of pre-training data, I would say it's pretty hard to do science on this type of question because you can't create that many pre-trained models. Maybe you can't train a GPT-4 sized model and do ablation studies at that scale. Maybe you can train a ton of GPT-2 size models or even a GPT-3 size model with different data blends and see what you get. I'm not aware of any public results on ablations involving code data and reasoning performance and so forth. I'd be very interested to know about those results.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.