High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“Because of the ability to do pipelining, a rack at least should not be a constraint on your ability to fit the model parameters.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…This goes back to the promise at the beginning of the lecture, which was this will actually tell you about AI progress as well. To the extent it is the case that model size scaling has been slow until recently… Let me make sure I understand the claim. The claim would not be you could have trained across more racks. It was just that it would not have made sense before, we didn’t have the ability to do inference for a bigger model easily. Actually, pipelining doesn’t help with context length. It totally helps with model size. Because of the ability to do pipelining, a rack at least should not be a constraint on your ability to fit the model parameters. The other consideration you’re asking is, why hasn’t it scaled up more, and why did bigger scale-up domains help? We talked through one aspect of that, which is that it’s not because of memory capacity. We have a solution to the memory capacity at least with respect to model size, not with respect to KV cache size but at least with respect to model size. The other issue that shows up is latency. I was just about to ask, going from rack to rack, what is the latency cost per hop?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence