Evidence receipt / evaluation
Published · transcript-backedDwarkesh Patel: evaluation
15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
“Correct me if this is wrong. It seems like you’re implying that right now we have models that are on a per token basis pretty smart.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…At a high level, I would agree that longer-horizon tasks are going to require more model intelligence to do well. They are going to be more expensive to train. I'm not sure I would expect a really clean scaling law unless you set it up in a very careful way, or design the experiment in a certain way. There might end up being some phase transitions where once you get to a certain level you can deal with much longer tasks. For example, when people do planning for different timescales, I'm not sure they use completely different mechanisms. We probably use the same mental machinery thinking about one month from now, one year from now, or a hundred years from now. We're not actually doing some kind of reinforcement learning where we need to worry about a discount factor that covers that timescale and so forth. Using language, you can describe all of these different timescales and then you can do things like plan. In the moment you can try to make progress towards your goal, whether it's a month away or 10 years away. I don’t know if it’s a phase transition but I might expect the same out of models where there might be some capabilities that work at multiple scales. Correct me if this is wrong. It seems like you’re implying that right now we have models that are on a per token basis pretty smart. They might be as smart as the smartest humans on a per token basis. The thing that prevents them from being as useful as they could be is that five minutes from now, they're not going to be still writing your code in a way that’s coherent and aligns with your broader goals you have for your project or something. If it's the case that once you start this long-horizon RL training regime it immediately unlocks your ability to be coherent for longer periods of time, should we be predicting something that is human-level as soon as that regime is unlocked? If not, then what is remaining after we can plan for a year and execute projects that take that long? It's not totally clear what we're going to see once we get into that regime or how fast progress will be. That's still uncertain. I wouldn't expect everything to be immediately solved by doing any training like this. There'll be other miscellaneous deficits that the models have that cause them to get stuck or make worse decisions than humans. I don’t expect that this one little thing will unlock all capabilities. But some improvement in the ability to do long-horizon tasks might go quite far.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.