High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“Sorry, I think the way I said it was super garbled. Just for the audience, forward plus backwards per parameter is 6.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…Somewhere in the range of two to six. We’ll say somewhere in the range of two to six and leave it at that. Then we can add in the inference cost. The inference cost is two, the number of active parameters times the data in inference. Sorry, I think the way I said it was super garbled. Just for the audience, forward plus backwards per parameter is 6. Forward alone is 2. That’s why RL, where you’re definitely going to generate all the trajectories but you might or might not train all the trajectories, is 2 to 6. Yes. Thank you. And then inference is just 2. We’re going to solve for essentially equality of all three of these terms. That is the ballpark of where people are going to be. Labs have more information on what is productive in doing more RL, for example, versus doing more pre-training. I don’t have that information, but I think a good ballpark is a 33% split between each of them.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence