High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Dwarkesh Podcast / episode intelligence

Reiner Pope – The math behind how LLMs are trained and served

29 Apr 2026 24 published claims 2 attributable people

Speakers in the public record

Claim mix

evaluation 10belief 10commitment 2recommendation 1preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

24 published records

01 / evaluation

Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say.

“Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

02 / evaluation

It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that.

“It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

17 / evaluation

I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights.

“I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

22 / evaluation

Oh, interesting. This is sort of obvious, but the difference between micro-batch and batch doesn’t matter at all in inference because you can just call it whatever you want.

“Oh, interesting. This is sort of obvious, but the difference between micro-batch and batch doesn’t matter at all in inference because you can just call it whatever you want.”
Speaker
Dwarkesh Patel
Publisher
Dwarkesh Podcast

23 / evaluation

If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds.

“If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

24 / evaluation

A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU.

“A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast
Search evidence