High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

John Collison: evaluation

26 Feb 2026 Cheeky Pint Reiner Pope of MatX on accelerating AI with transformer-optimized chips

“I don't know if they've talked about this stuff publicly, but you know that their internal metrics similarly show that the faster the product page loads, the more people buy it.”

— John Collison

Source trail

Everything needed to verify it.

Speaker
John Collison
Attribution
Verified speaker
Claim type
evaluation
Recorded
26 Feb 2026
Publisher
Cheeky Pint

Transcript context

…Something like that. In result times where they noticed more Google engagement, the faster the results were, and you'd think that 50 milliseconds is imperceptible to a human. It almost is, but turns out it's not. I think Amazon has, certainly, they've optimized the latency of the Amazon experience quite a lot. I don't know if they've talked about this stuff publicly, but you know that their internal metrics similarly show that the faster the product page loads, the more people buy it. Yet in AI, Google has carved out a meaningful advantage via Gemini just being really fast for its level of intelligence. As far as I can tell, ahead of most of the other labs at a latency, at a fixed high level of intelligence. Why have you guys or Groq or better chips not been adopted faster to give this product latency? It's just that this will happen and you guys will be powering all the AI products, but I note that Google has an interesting lead there. I think there's ultimately… At least for existing chips in the market, there's a really uncomfortable trade-off between latency and throughput. The chips that are best at throughput have historically been the chips that are based on HBM as the memory. That is Google, Amazon, NVIDIA. In order to have very large throughput, you need a lot of inferences in flight simultaneously. That needs a large memory. But that hasn't been so good at latency. Then, there's the Groq and Cerebras that are much better at latency because they've got this, the SRAM, weights are in SRAM, very low latency. The problem is, and the challenge when you go to a Groq or Cerebras system is that the throughput you get there, it just is not very good. The fundamental dollars per token is just not competitive with Google or NVIDIA or Amazon. It is actually possible to do both in the same chip. It's kind of an obvious thing. You say you take the HBM, you take the SRAM, put them together on the same chip, you put the weights in SRAM, and you put all of the inference data in HBM. That is what we are doing, in fact. I think that actually hits a really nice sweet spot where you can get low latency and also be very cheap. I think that's a really attractive point to be. It hasn't happened in the market yet, just because of product decisions that have been made by the different chips.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence