Evidence receipt / observation
Published · transcript-backedDwarkesh Patel: observation
4 Jun 2024 Dwarkesh Podcast Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history
“On that point in particular, many people who have longer timelines have come on the podcast and made the point that the way to train this long horizon RL, it's not… Earlier we were talking about how they can think for five minutes, but not for longer.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 4 Jun 2024
- Publisher
- Dwarkesh Podcast
Transcript context
…We'll see. On sample efficiency, it’s sort of first principles but there's this clear missing middle. People hadn't been trying. Now people are really trying. Again, often in deep learning something like the obvious thing works and there are a lot of details to get right. It might take some time, but now people are really trying. We will get a lot of signal in the next couple of years on unhobbling. What is the signal on unhobbling that would be interesting? The question is basically, are you making progress on test time compute? Is this thing able to think longer horizon than just a couple hundred tokens? That was unlocked by chain-of-thought. On that point in particular, many people who have longer timelines have come on the podcast and made the point that the way to train this long horizon RL, it's not… Earlier we were talking about how they can think for five minutes, but not for longer. It's not because they can't physically output an hour's worth of tokens. Even Gemini has a million in context, and the million of context is actually great for consumption. It solves one important hobbling, which is the onboarding problem. A new coworker in your first five minutes, like a new smart high school intern, is not useful at all. A month in, they’re much more useful because they've looked at the monorepo, understand how the code works, and they've read your internal docs. Being able to put that in context solves this onboarding problem. They're not good at the production of a million tokens yet.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.