High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Reiner Pope: recommendation

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“If you were designing your hardware, you could say, “I didn’t need that much memory because I don’t need the weights to fit in one rack.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
recommendation
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…Makes a ton of sense. So it’s a no-brainer to use pipelining during inference, but there’s this harder trade-off during training. Even in inference, in fact, it is not used a ton. It reduces your memory capacity requirements, but there’s actually a huge surplus. I think you were saying that a rack of Blackwell has many tens of terabytes. That’s much bigger than a trillion parameter model. A trillion parameter model only needs one terabyte, so it already fits. There’s not a huge benefit from pipelining because you’re reducing a number that’s already pretty small. But it does say that theoretically, maybe you had too much memory there. You could have built different hardware that has less memory. If you were designing your hardware, you could say, “I didn’t need that much memory because I don’t need the weights to fit in one rack. I can fit the weights in eight racks, then I could have built hardware that didn’t have so much HBM per GPU.” Macro question: everybody’s talking about the memory wall right now. Memory is getting super expensive. There’s not enough memory. Smartphone volume will go down 30% because there’s not enough memory. This is shocking, Dylan said hyperscalers are spending 50% of their CapEx this year on memory.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence