High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

John Schulman: belief

15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

“On the other hand, I would also expect to gain a lot from doing practice at training time. So I think that you’d get the best results by combining these two things.”

— John Schulman

Source trail

Everything needed to verify it.

Speaker
John Schulman
Attribution
Verified speaker
Claim type
belief
Recorded
15 May 2024
Publisher
Dwarkesh Podcast

Transcript context

…I’ve heard two ideas in terms of using that internal monologue to get better at reasoning. At least publicly, I've seen two ideas and I'm curious which one you think is more promising. One is that the model learns from its outputs over a bunch of potential trains of thought, and it learns to follow the one that leads to the correct answer. It is then trained on that before deployment. The other one is you use a bunch of compute to do inference in deployment. This approach involves the model talking to itself while it's deployed. Which one do you expect to be closer to the way a model has been trained when it gets really good at reasoning? Is it because it's doing just a bunch of inference clouds? Is it just because you've trained it to do well at that? You could define reasoning as tasks that require some kind of computation at test time or maybe some kind of deduction. By definition, reasoning would be tasks that require some test time computation and step-by-step computation. On the other hand, I would also expect to gain a lot from doing practice at training time. So I think that you’d get the best results by combining these two things. Right now, you have these two ways the model learns. One is in training, whether it's pre-training or post-training. Most of the compute in training is spent on pre-training, glossing over trillions of tokens, skimming trillions of tokens worth of information. If a human was subjected to that, they would just be totally confused. It's just not a very efficient way to learn. The other way is in-context learning. Of course that is more sample-efficient, but it's destroyed with each instance. I'm curious if you think that there's a path for something in between those, where it’s not destroyed at each instance but it's also not as frivolous as just seeing trillions of tokens. Something more deliberate and active.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence