Evidence receipt / evaluation
Published · transcript-backedJohn Schulman: evaluation
15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
“There are some algorithms that work this way, like mixture models or multiplicative weight update algorithms, where you have—I don’t want to say mixture of experts because it means something different—basically a weighted combination of experts with some learned gating.”
Source trail
Everything needed to verify it.
- Speaker
- John Schulman
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…This might have a very simple answer. You train bigger models on the same amount of data and they become smarter. Or to get the same level of intelligence, you only have to train them on less data. Why is that the case? It's got more parameters, seen fewer things, and now it's equally as smart. Why is that? I don't think anyone has a good explanation for the scaling law with parameter count. I don't even know what the best mental model is for this. Clearly, you have more capacity if you have a bigger model. So you should eventually be able to get lower loss. Why are bigger models more sample efficient? I can give you a sketchy explanation. You could say that the model is an ensemble of different circuits that do the computation. You could imagine that it's doing computations in parallel and the output is a weighted combination of them. If you have more width… actually width is somewhat similar to depth because with residual networks, depth can do something similar to width in terms of updating what's in the residual stream. You're learning all these different computations in parallel and you have more of them with a bigger model. So you have a higher chance that one of them is lucky, ends up guessing correctly a lot, and gets upweighted. There are some algorithms that work this way, like mixture models or multiplicative weight update algorithms, where you have—I don’t want to say mixture of experts because it means something different—basically a weighted combination of experts with some learned gating. I actually said something slightly wrong, but you could imagine something like that. Just having a bigger model gives you more chances to get the right function. Of course, it's not just totally disjoint functions you're taking a linear combination of. It's more like a library where you might chain the functions together in some way. There's some composability. So I would say a bigger model has a bigger library of different computations, including lots of stuff that's dormant and only being used some of the time, but it has more space to look for circuits to do something useful. Stepping back from the current research questions, I want to understand your modal scenario of what happens for the next few years. Towards the beginning of the conversation, we were talking about the case in which it progresses really fast, but let's just take the modal scenario. You're unlocking long-horizon RL at some point, but as you said, there are potentially other bottlenecks. What's happening? How good are these models? How are they being deployed? What other modalities are part of them and at what stage are these being unlocked? I want to understand your broader picture of what the next few years look like.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.