High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…Got it. Is sparse attention what everybody uses in practice? I’m pretty excited about sparse attention. It’s hard to know what the labs are using. DeepSeek has published a sparse attention mechanism. I’ll just put a plug in that some of the DeepSeek papers that have published sparse attention end up putting a square root in this term. So far, we’ve looked at the latency. It’s hard to read off cost from this. If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds. I have to pay the rental time for that time. So it’s $2/hour per GPU or something like that. That’s the cost of this inference, but how many tokens have I processed during that inference? That is the batch size. What we actually want to plot is the cost versus batch size, which is t over B versus batch size. This is the cost per token. We have to imagine dividing each of these three curves by B, so multiplying by this reciprocal. What we end up with there is… The compute curve was linear. We divide by B, and that makes it a constant here. This is t compute. The KV fetch was linear, and now it becomes a constant as well. Then the weight fetch was constant, and now we’ve divided by B, so it becomes this hyperbola. Again, we’re going to compute the max of the sum. The sum of these two terms shifts the hyperbola up. The sum of the KV fetch and the weight fetch gives us a higher hyperbola that’s like this. Then we’re going to take the max with the compute here. We end up with this being the overall shape that we care about. Again, we see some limiting behavior. The cost initially starts very high at a batch size of one. It almost goes to infinity because we’ve got so many weight fetches that are not amortized over a large batch size. But as we increase the batch size, the weight fetches become amortized over so many different batch elements that their cost grows very small, and eventually the compute time ends up driving the cost. So there is a limiting lower bound on cost, which is this line here. So Claude Code Slow or Codex Slow or whatever would just live on this line. It wouldn’t help much because you’re not able to amortize the KV values over a much bigger batch.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence