Evidence receipt / belief
Published · transcript-backedReiner Pope: belief
29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served
“If we think that the memory in a rack is a bottleneck, then there’s a constraint on how fast we can go.”
Source trail
Everything needed to verify it.
- Speaker
- Reiner Pope
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 29 Apr 2026
- Publisher
- Dwarkesh Podcast
Transcript context
…So there’s a talk by Ilya where he says, “Today we know not to do pipeline parallelism.” And Horace He gave my friends and me… I hate that it sounds like a Dr. Seuss quote. But he gave us a lecture on these different kinds of parallelisms. He said the problem with pipeline parallelism is that, other than the bubbles, it creates these architectural constraints. Kimi, for example, has these residuals where attention attends to layers a few back, so it becomes hard to implement in this way. I guess we didn’t fully articulate even what is the benefit that we’re getting from pipelining. These complexities are real. Pipelining is a massive hassle, but it does give you some benefits. You can then decide whether those benefits are worth the costs. It has some benefits in inference, maybe bigger benefits in training. In inference, what are we saving on? Are we saving on memory time or compute time? Not really. We’re just moving the memory time from one chip to another chip, or one rack to a different rack. There’s no actual benefit in runtime. However, what we are saving on is memory capacity. If we think that the memory in a rack is a bottleneck, then there’s a constraint on how fast we can go. Pipelining allows us to massively reduce that bottleneck. The opposite connotation to this… Before this interview, I was chatting with Axel, who’s a GPU performance engineer at Jane Street. He was explaining that to do pipelining, you have to do micro-batches rather than full batches. If you do micro-batches, then you’re by definition not able to amortize loading the weights across all the users or all the sequences. The positive connotation of that is you don’t have to use as much memory. The negative connotation is that we can’t amortize loading the weights across all those users. Maybe it’s worth explaining why you have to do micro-batches.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.