Evidence receipt / belief
Published · transcript-backedJohn Schulman: belief
15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
“There are a lot of different, separate axes for improvement. We think about data quality, data quantity.”
Source trail
Everything needed to verify it.
- Speaker
- John Schulman
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 15 May 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…The current GPT-4 has an Elo score that is like a hundred points higher than the original one that was released. Is that all because of what you're talking about, with these improvements that are brought on by post-training? Yeah, most of that is post-training. There are a lot of different, separate axes for improvement. We think about data quality, data quantity. There’s just doing more iterations of the whole process of deploying and collecting new data. There’s also changing what kind of annotations you're collecting. There's a lot of things that stack up but together they give you a pretty good effective compute increase. That's a huge increase. It's really interesting that there's this much room for improvement from post-training. What makes for somebody who's really good at doing this sort of RL research? I hear it's super finicky. What is the sort of intuition that you have that enables you to find these ways to mess with the data and set up these environments?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.