Speakers in the public record
Claim mix
evaluation 10belief 10commitment 2recommendation 1preference 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
24 published records
“Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say.”
- Publisher
- Dwarkesh Podcast
“It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that.”
- Publisher
- Dwarkesh Podcast
“Ok, so I was asking you, have neural networks actually been used for cryptography? And we realized it may be better to just do this on the blackboard.”
- Publisher
- Dwarkesh Podcast
“I don’t have that information, but I think a good ballpark is a 33% split between each of them.”
- Publisher
- Dwarkesh Podcast
“If you look at Hopper, you had eight Hoppers, and I think that’s 640 gigabytes as of 2022.”
- Publisher
- Dwarkesh Podcast
“I think that’s a pretty good assumption. If you look at the numbers, it might also turn out that it’s one tier down, and it’s DDR versus flash.”
- Publisher
- Dwarkesh Podcast
“I can motivate the batch at least a little bit. We will see exactly why batch is such a favorable optimization.”
- Publisher
- Dwarkesh Podcast
“I think if you do it with a spreadsheet and actually model it out, you might notice when the money’s going down the drain.”
- Publisher
- Dwarkesh Podcast
“I think often the number of active parameters could be in the range of a hundred billion, something like that.”
- Publisher
- Dwarkesh Podcast
“I think starting with equalizing in cost is right, but depending on how you model the cost, this comes close to equalizing in data.”
- Publisher
- Dwarkesh Podcast
“I think we just have to make some real-world assumptions here in order to do that.”
- Publisher
- Dwarkesh Podcast
“We will directly jump to the ratio of the time spent on scale-up over the time spent on scale-out.”
- Publisher
- Dwarkesh Podcast
“Sorry, I think the way I said it was super garbled. Just for the audience, forward plus backwards per parameter is 6.”
- Publisher
- Dwarkesh Podcast
“If we think that the memory in a rack is a bottleneck, then there’s a constraint on how fast we can go.”
- Publisher
- Dwarkesh Podcast
“If you were designing your hardware, you could say, “I didn’t need that much memory because I don’t need the weights to fit in one rack.”
- Publisher
- Dwarkesh Podcast
“I think that indicates that this is the reasonably balanced cost point, and going massively beyond that would be cost-prohibitive.”
- Publisher
- Dwarkesh Podcast
“I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights.”
- Publisher
- Dwarkesh Podcast
“I’m kind of shocked to see spinning disk being used at all, because it’s such an old technology.”
- Publisher
- Dwarkesh Podcast
“Because of the ability to do pipelining, a rack at least should not be a constraint on your ability to fit the model parameters.”
- Publisher
- Dwarkesh Podcast
“If I think about what that means, it means I actually have time to read the HBM twice.”
- Publisher
- Dwarkesh Podcast
“What that means is, given that I know I’m going to be holding something for five minutes, I would like to pick a memory that I can read every five minutes.”
- Publisher
- Dwarkesh Podcast
“Oh, interesting. This is sort of obvious, but the difference between micro-batch and batch doesn’t matter at all in inference because you can just call it whatever you want.”
- Publisher
- Dwarkesh Podcast
“If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds.”
- Publisher
- Dwarkesh Podcast
“A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU.”
- Publisher
- Dwarkesh Podcast