High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“I think that indicates that this is the reasonably balanced cost point, and going massively beyond that would be cost-prohibitive.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…There’s this idea that Dario said on the podcast, and others have said, which is, “We don’t need continual learning for AGI, in-context learning is enough.” If you believe that, then you have to think that we have to get to a hundred-million-token context length to have an employee that is the equivalent of working with you for a month. Now, maybe that’s no longer true with sparse attention or something. But if you think that, then some ML infra thing would have to change to allow for a hundred million, like the memory bandwidth, to allow for a hundred-million-token context lengths. Sparse attention gives you a get-out for sure, because you get this square root. It gives you a big improvement. But if you look at the history of context lengths of models, from earlier models like GPT-3, maybe to GPT-4—I don’t remember when the transition happened exactly—they shot up from about 8K to 100-200K. And then for the last year or two, they’ve all been hovering around there. I think that indicates that this is the reasonably balanced cost point, and going massively beyond that would be cost-prohibitive. Not because of the compute cost, because of the memory bandwidth...…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence