Evidence receipt / belief
Published · transcript-backedLex Fridman: belief
3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
“I think the pricing model mostly they use for input tokens is about one fourth of price of the output tokens.”
Source trail
Everything needed to verify it.
- Speaker
- Lex Fridman
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Feb 2025
- Publisher
- Lex Fridman Podcast
Transcript context
…I can explain that. Today, if you use a model, like you look at an API, OpenAI charges a certain price per million tokens. And that price for input and output tokens is different. And the reason is is that when you’re inputting a query into the model, let’s say you have a book, that book, you must now calculate the entire KV cache for this, key-value cache. And so when you do that, that is a parallel operation. All of the tokens can be processed at one time and therefore you can dramatically reduce how much you’re spending. The FLOP requirements for generating a token and an input token are identical. If I input one token or if I generate one token, it’s completely identical. I have to go through the model. But the difference is that I can do that input, i.e. the prefill, i.e. the prompt, simultaneously in a batch nature and therefore it is all FLOP. I think the pricing model mostly they use for input tokens is about one fourth of price of the output tokens. Correct. But then output tokens, the reason why it’s so expensive is because I can’t do it in parallel. It’s autoregressive. Every time I generate a token, I must not only read the whole entire model into memory and activate it, calculate it to generate the next token, I also have to read the entire KV cache. And I generate a token and then I append that one token I generated and it’s KV cache and then I do it again. And so therefore, this is a non-parallel operation. And this is one where you have to, in the case of prefill or prompt, you pull the whole model in and you calculate 20,000 tokens at once, 20,000-…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.