High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Reiner Pope: commitment

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“We will directly jump to the ratio of the time spent on scale-up over the time spent on scale-out.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
commitment
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…But the communication within the rack is batch size times number of GPUs. Number of activated GPUs. I don’t send to this GPU at all. There’s an explosion from 1-3x larger here in this diagram. The key thing is that I didn’t even need to send to this GPU at all, and so that’s a big saving. We’re going to talk through to what extent scale-up is a bottleneck over scale-out. We will directly jump to the ratio of the time spent on scale-up over the time spent on scale-out. This is the quantity we’re talking about. The first consideration is that scale-up is 8x faster than scale-out generally. At a baseline, if the bandwidths were the same, we would have this 1/8, which is coming from bandwidth. But then we have some amount of expansion in how much data we’re sending. If one token comes in here, then this one token gets routed to, in the DeepSeek case maybe 32 experts or 16 experts. It gets routed to some number of experts. So this is the number of activated experts. This same thing applies on multiple different layers, so maybe I’m going to run two layers. There’s also multiple times the number of layers per stage. Don’t you need to multiply the whole thing by two for the all-to-all?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence