High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Shawn Wang: belief

8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

“Let’s pick on hardware efficiency since you also have, you also track hardware stuff. And I think the general assertion or the message is that the efficiency from next gen Nvidia chips is actually not 4X.”

— Shawn Wang

Source trail

Everything needed to verify it.

Speaker
Shawn Wang
Attribution
Verified speaker
Claim type
belief
Recorded
8 Jan 2026
Publisher
Latent Space

Transcript context

…I think that’s right. There’s a number of drivers at play and we kind of outline kind of. Six key ones here. Um, but you know, as complex as changing quickly, all of these have changed very dramatically in the last, uh, in the last 12 months. Let’s pick on hardware efficiency since you also have, you also track hardware stuff. And I think the general assertion or the message is that the efficiency from next gen Nvidia chips is actually not 4X. So you have what? 3X or 4X? You have 3X in here and it’s, it’s like 2X maybe, or it’s more of like a. Power story rather than like a share sort of compute tokens efficiency story. But yeah, what, what’s going on in, in hardware. Okay. So the, the, the, the odds, unfortunately, uh, is it depends and it just depends massively on like so many things across a bunch of different types of workloads and ways to think about it. So one of the simplest ways to think about this is to take single relevant model, to think about serving it at speeds that are realistic for what you actually might want to hit. And can afford to hit, and then think about the throughput per GPU that you can achieve serving the model at those speeds. Rease. One of the reasons that’s important is that there’s a trade-off between the throughput per GPU that you can achieve and the per user speed that you can achieve. And as a, it costs more to serve stuff fast to, to users. When you run all of that for especially big sparse models, you can get a lot better than two or three X gain going from Hopper to Blackwell generation to video. I am. This shouldn’t be too controversial. Let’s say I’m like, I’m. I’m pretty confident that Blackwell has delivered pretty enormous gains and that the next couple of years of NVIDIA’s roadmap are going to continue to deliver quite enormous gains and that those will actually come through as lower total cost per token to the companies that are running models on them and will allow bigger models will allow way more tokens to be made for lower cost and that that’s gonna continue these things also stack on all of the software and model improvements. So basically like my prediction across like both sides of that, like smile chart, uh, that we’re gonna see the left-hand side continue to be true and probably like for another order of magnitude and the right-hand side continue to be true for another order of magnitude, and that’s gonna enable a whole lot of things.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence