Speakers in the public record
Claim mix
evaluation 7preference 2belief 1uncertainty 1recommendation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
12 published records
“If I think about how the brain works versus what you’re describing here, at a high level the differences might be that while you can do structured sparsity in these accelerators and save yourself some area that you would have otherwise had to dedicate to gates, in the brain there’s unstructured sparsity.”
- Publisher
- Dwarkesh Podcast
“I don’t know if you have any commentary on what the brain might be doing versus how these chips work.”
- Publisher
- Dwarkesh Podcast
“It hurts your throughput, in fact, because the throughput of your chip is the product of how much you get done per clock cycle—which is based on area efficiency—times how many clocks you get per second.”
- Publisher
- Dwarkesh Podcast
“The trade-off is that the first FPGA costs you $10,000, whereas the first ASIC you make costs $30 million because it requires an entire tape-out.”
- Publisher
- Dwarkesh Podcast
“Instead of forming a running sum of everything that comes in, I will actually have two different running sums.”
- Publisher
- Dwarkesh Podcast
“I have this cost, which is 3 x n x p AND gates over here, compared to p x q gates in the actual circuit that is doing the thing I care about.”
- Publisher
- Dwarkesh Podcast
“We don’t want to bring this full matrix in from the register file every cycle, because that would cost too much in terms of wiring from the register file.”
- Publisher
- Dwarkesh Podcast
“You’re saying that since we’re only going to be loading this in once, let’s minimize bandwidth, because bandwidth equals die area.”
- Publisher
- Dwarkesh Podcast
“We’ve talked about how it seems obvious that you should try to maximize compute relative to communication.”
- Publisher
- Dwarkesh Podcast
“Really there’s only one addition per dot product, but we like to start with zero.”
- Publisher
- Dwarkesh Podcast
“If you split it in the middle, you can hit twice the clock frequency. That’s great, you get twice the performance, but at the cost of an extra register, which means more storage.”
- Publisher
- Dwarkesh Podcast
“I think the big observation you’ve made is that there’s this quadratic scaling with bit width, which is very effective and is the single reason low-precision arithmetic has worked so well for neural nets.”
- Publisher
- Dwarkesh Podcast