High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

22 May 2026 Dwarkesh Podcast Reiner Pope – Chip design from the bottom up

“I think the big observation you’ve made is that there’s this quadratic scaling with bit width, which is very effective and is the single reason low-precision arithmetic has worked so well for neural nets.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
22 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…Though it should be 4x. Yeah. What I’ve shown here is the simplest case of integer multiply. When you’re dealing with floating point, as you do in FP4 and FP8, there’s this other term, the exponent, that complicates the calculation. What can we see already from this? I think the big observation you’ve made is that there’s this quadratic scaling with bit width, which is very effective and is the single reason low-precision arithmetic has worked so well for neural nets. The other thing we’re going to do now is compare the area spent on the multiplication itself with all the circuitry around it. We’ll walk back in time a little bit and see how GPUs prior to Tensor Cores worked, which is in fact the same way CPUs worked. Where do we stick this multiply-accumulate unit? Generically, I’ll describe a CUDA core or a CPU. You’ll have some register file which stores some number of entries, maybe eight entries of, in this case, 4-bit numbers, but typically 32-bit numbers. Inside the CUDA core, I’ll have some register file of some depth, and then I’ll have my multiply-accumulate circuit. What it’s going to do is take three arbitrary registers from this register file, perform the multiply-accumulate, and then write back to the register file. It’s going to write to this one, but it was able to read from this one, this one, and another random one. It will take three inputs like this. This is the core data path of many processors. Most processors look like this. You’ve got some set of registers, and then you’ve got some set of logic units, or ALUs. We want to analyze the cost of the data movement from the register file to the ALU and back. Ultimately, there’s going to be some circuit that says, “Well, I don’t always have to select this guy. I might select any of the registers at any point in time.” The first question is: how can I build a circuit? The circuit I’m going to look for is a mux. In this case, it’s going to have eight inputs, one from each entry of the register file, and it’s going to have one output, which is actually producing this output. What is the cost of this thing? All we have to build it out of is AND and OR. How do we build it? We do the dumbest thing possible. We form a mask. When we want to read the third entry, we’re going to AND every single entry with either 1 or 0 based on whether that’s what we want to read, and then we’re going to OR all of them together. Just to make sure I understand the basics. What the mux is doing is just selecting an input?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence