High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“If you look at Hopper, you had eight Hoppers, and I think that’s 640 gigabytes as of 2022.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…2023. Okay. And it was rumored to be over one trillion parameters. It seems like only now, within the last six months, have models been getting released that have significantly more parameters than the model released three years ago, when supposedly there should have been this scaling in the meantime. Is the reason that we were just waiting for racks with enough memory to hold a five-trillion parameter model, along with its KV cache for enough users for a lot of sequences? Or if you’re doing RL, a similar consideration of actually holding the KV cache for the batch of problems you’re trying to solve. If you look at Hopper, you had eight Hoppers, and I think that’s 640 gigabytes as of 2022. With Blackwell finally, which was deployed in…? Very recently. Maybe last year.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence