High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Dwarkesh Patel: observation

13 Mar 2026 Dwarkesh Podcast Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI compute

“I heard a theory that the reason is that Nvidia’s scale-ups have just not had that much memory capacity.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
observation
Recorded
13 Mar 2026
Publisher
Dwarkesh Podcast

Transcript context

…Earlier I was mentioning how communication within a chip is super fast. Communication within chips that are in the same rack is fast, but not as fast. It’s on the order of terabytes. Communication very far away is on the order of hundreds of gigabytes. As you get further distance, maybe across the country, the order of magnitude is on the order of gigabytes. A scale-up domain is this tight domain where the chips are communicating on the order of terabytes a second. For Nvidia, previously this meant an H100 server had eight GPUs, and those eight GPUs could talk to each other at terabytes a second. With Blackwell NVL72, they implemented rack-scale scale-up. That meant all seventy-two GPUs in the rack could connect to each other at terabytes a second. The speed doubled generation on generation, but the most important innovation was going from eight to seventy-two in the domain. When we look at Google, their scale-up domain is completely different. It has always been on the order of thousands. With TPU v4, they had pods the size of four thousand chips. With v8 or v7, they have pods in the eight or nine thousand range. What’s relevant here is that it’s not the same as Nvidia. It’s not like for like. Google has a topology that’s a torus. Every chip connects to six neighbors. Nvidia’s 72 GPUs connect all-to-all. They can send terabytes a second to any arbitrary other chip in that pod of scale-up. Whereas Google, you have to bounce through chips. If TPU 1 needs to talk to TPU 76, it has to bounce through various chips, and there is always some blocking of resources when you do that because that one TPU is only connected to six other TPUs. So there is a difference in topology and bandwidth, and there are trade-offs and advantages to both. Google gets to have a massive scale-up domain, but they have the trade-off of bouncing across chips to get from one to another. You can only talk to six direct neighbors. Amazon has mutated their scale-up domain. They’re somewhere in between Nvidia and Google. They’re trying to make larger scale-up domains. They try to do all-to-all to some extent with switches, which is what Nvidia does, but they also use torus topologies like Google to some extent. As we advance forward to next generations, all three of them are moving more towards a dragonfly topology. That means there are some fully connected elements and some elements that are not fully connected. You can get the scale-up to be hundreds or thousands of chips, but also have it not contend for resources when bouncing through chips. Related question: I heard somebody make the claim that the reason parameter scaling has been slow—and only now are we getting bigger models from OpenAI and Anthropic—is that… The original GPT-4 is over a trillion parameters, and only now are models starting to approach that again. I heard a theory that the reason is that Nvidia’s scale-ups have just not had that much memory capacity. Let’s say you have a 5T model running at FP8, so that’s five trillion gigabytes. And then you have the KV cache, let’s say it’s— Just call it the same size.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence