High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Reiner Pope: belief

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“I don’t have that information, but I think a good ballpark is a 33% split between each of them.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
belief
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…Sorry, I think the way I said it was super garbled. Just for the audience, forward plus backwards per parameter is 6. Forward alone is 2. That’s why RL, where you’re definitely going to generate all the trajectories but you might or might not train all the trajectories, is 2 to 6. Yes. Thank you. And then inference is just 2. We’re going to solve for essentially equality of all three of these terms. That is the ballpark of where people are going to be. Labs have more information on what is productive in doing more RL, for example, versus doing more pre-training. I don’t have that information, but I think a good ballpark is a 33% split between each of them. I’m not sure I understand the intuition for that. Another naive model could have been that RL plus pre-training would be 50% and inference would be 50%.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence