Evidence receipt / evaluation
Published · transcript-backedReiner Pope: evaluation
29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served
“Because of the ability to do pipelining, a rack at least should not be a constraint on your ability to fit the model parameters.”
Source trail
Everything needed to verify it.
- Speaker
- Reiner Pope
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 29 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…This goes back to the promise at the beginning of the lecture, which was this will actually tell you about AI progress as well. To the extent it is the case that model size scaling has been slow until recently… Let me make sure I understand the claim. The claim would not be you could have trained across more racks. It was just that it would not have made sense before, we didn’t have the ability to do inference for a bigger model easily. Actually, pipelining doesn’t help with context length. It totally helps with model size. Because of the ability to do pipelining, a rack at least should not be a constraint on your ability to fit the model parameters. The other consideration you’re asking is, why hasn’t it scaled up more, and why did bigger scale-up domains help? We talked through one aspect of that, which is that it’s not because of memory capacity. We have a solution to the memory capacity at least with respect to model size, not with respect to KV cache size but at least with respect to model size. The other issue that shows up is latency. I was just about to ask, going from rack to rack, what is the latency cost per hop?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.