High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Lambert: belief

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model.”

— Nathan Lambert

Source trail

Everything needed to verify it.

Speaker
Nathan Lambert
Attribution
Verified speaker
Claim type
belief
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…And those innovations are going to help you with … I mean, your highly memory constrained in this? They help with memory constraint and performance. Gemini is the model that has the longest context length that people are using. Gemini is known for one million and now two million context length. You put a whole book into Gemini and sometimes it’ll draw facts out of it. It’s not perfect, they’re getting better. So there’s two things. It’s, one, to be able to serve this on the memory level. Google has magic with their TPU stack where they can serve really long contexts. And then there’s also many decisions along the way to actually make long context performance work that supplies the data. There’s subtle changes to these computations in attention and it changes the architecture. But serving long context is extremely memory constrained, especially when you’re making a lot of predictions. I actually don’t know why input and output tokens are more expensive, but I think essentially output tokens, you have to do more computation because you have to sample from the model. I can explain that. Today, if you use a model, like you look at an API, OpenAI charges a certain price per million tokens. And that price for input and output tokens is different. And the reason is is that when you’re inputting a query into the model, let’s say you have a book, that book, you must now calculate the entire KV cache for this, key-value cache. And so when you do that, that is a parallel operation. All of the tokens can be processed at one time and therefore you can dramatically reduce how much you’re spending. The FLOP requirements for generating a token and an input token are identical. If I input one token or if I generate one token, it’s completely identical. I have to go through the model. But the difference is that I can do that input, i.e. the prefill, i.e. the prompt, simultaneously in a batch nature and therefore it is all FLOP.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence