High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…This brings us back to the question then, is the size of the scale-up at all relevant to why AI model sizes have been what they have been over the last few years, whether through training or through inference? We talked about latency of the hop. There is also just the tmem latency. The memory time latency is actually massively improved by larger scale-up domains. I’ll recall tmem down here. tmem for the weights was equal to the number of total parameters divided by the memory bandwidth. Which memory bandwidth are we talking about here? Is it just one GPU? It is the number of GPUs that I can use in parallel to load these weights. I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights. This is actually extremely effective. Basically, I end up with a term here, this memory bandwidth term itself is equal to scale-up size... Times memory bandwidth per GPU.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence