Evidence receipt / prediction
Published · transcript-backedKyle Kranen: prediction
10 Mar 2026 Latent Space NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light" — Nader Khalil (Brev), Kyle Kranen (Dynamo)
“On the decode side because you’re doing a full Passover, all the weights and the entire sequence, every time you do a decode step and you’re, you don’t have the quadratic computation of KV cache, it’s usually memory bound because you’re retrieving a linear amount of memory and you’re doing a linear amount of compute as opposed to prefill where you retrieve a linear amount of memory and then use a quadratic.”
Source trail
Everything needed to verify it.
- Speaker
- Kyle Kranen
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 10 Mar 2026
- Publisher
- Latent Space
Transcript context
…gives you the knobs. Yeah, exactly. It gives you the knobs. Uh, and one thing that like we, we use a lot in contemporary inference and is, you know, starting to like pick up from, you know, in, in general knowledge is this co concept of disaggregation. So historically. Models would be hosted with a single inference engine. And that inference engine [00:39:00] would ping pong between two phases. There’s prefill where you’re reading the sequence generating KV cache, which is basically just a set of vectors that represent the sequence. And then using that KV cache to generate new tokens, which is called Decode. And some brilliant researchers across multiple different papers essentially made the realization that if you separate these two phases, you actually gain some benefits. Those benefits are basically a you don’t have to worry about step synchronous scheduling. So the way that an inference engine works is you do one step and then you finish it, and then you schedule, you start scheduling the next step there. It’s not like fully asynchronous. And the problem with that is you would have, uh, essentially pre-fill and decode are, are actually very different in terms of both their resource requirements and their sometimes their runtime. So you would have like prefill that would like block decode steps because you, you’d still be pre-filing and you couldn’t schedule because you know the step has to end. So you remove that scheduling issue and then you also allow you, or you yourself, to like split the work into two different ki types of pools. So pre-fill typically, and, and this changes as, as model architecture changes. Pre-fill is, right now, compute bound most of the time with the sequence is sufficiently long. It’s compute bound. On the decode side because you’re doing a full Passover, all the weights and the entire sequence, every time you do a decode step and you’re, you don’t have the quadratic computation of KV cache, it’s usually memory bound because you’re retrieving a linear amount of memory and you’re doing a linear amount of compute as opposed to prefill where you retrieve a linear amount of memory and then use a quadratic. You know, it’s funny, someone exo Labs did a really cool demo where for the DGX Spark, which has a lot more compute, you can do the pre the compute hungry prefill on a DG X spark and then do the decode on a, on a Mac. Yeah. And so…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.