High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

John Schulman: belief

15 May 2024 Dwarkesh Podcast John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

“I agree that you'd probably also want to supplement that with some kind of fine-tuning.”

— John Schulman

Source trail

Everything needed to verify it.

Speaker
John Schulman
Attribution
Verified speaker
Claim type
belief
Recorded
15 May 2024
Publisher
Dwarkesh Podcast

Transcript context

…Yeah. That does feel like something that's missing from today's systems. People haven't really pushed too hard on this middle ground between large-scale training—where you produce a single snapshot model that's supposed to do everything like a deployed model—and on the other hand in-context learning. Part of that is that we've just been increasing context length so much that there hasn't been an incentive for it. If you can go to a hundred thousand or a million context, then that's actually quite a lot. It’s not actually the bottleneck in a lot of cases. I agree that you'd probably also want to supplement that with some kind of fine-tuning. The capabilities you get from fine-tuning and in-context learning are probably somewhat complementary. I’d expect us to want to build systems that do some online learning and also have some cognitive skills, like introspecting on their own knowledge and seeking out new knowledge that fills in the holes. Is this all happening at the same time? Is it just a new training regime where all these things can happen at once, whether it’s long-horizon or this kind of training? Are they separate or not? Is the model smart enough to both introspect and act on longer horizons so that you get adequate reward on the long-horizon tasks?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence