Evidence receipt / observation
Published · transcript-backedMicah-Hill Smith: observation
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“We don’t actually publish anything on this right now, but have tracked it a bunch internally in our internal analytics on evals across all the models that we run, where we look at the difficulty to questions and the correlation between token usage and difficulty and net net, surprise, surprise, like models have got.”
Source trail
Everything needed to verify it.
- Speaker
- Micah-Hill Smith
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…in a non-reasoning model for the same task. That’s true. I think 5.1 was it. And then 5.1 Codex had these chart, which was super nice of this, like, let’s say bottom 10 percentile query being faster, but top 10 percentile being longer. And that’s a kind of the efficiency chart you want to see, right? Yeah. So that is an extra thing. Let’s say that we’ve got, that’s a really important extra thing though, right? That you’ve got not just the average number of token span used by the model, which we cover really well right now, but the behavior that you want in the model is it to use more tokens when it needs more tokens and not to use more tokens when it doesn’t need more tokens. So that’s what OpenAI, we’re basically claiming that 5.1 Codex is better at. We don’t actually publish anything on this right now, but have tracked it a bunch internally in our internal analytics on evals across all the models that we run, where we look at the difficulty to questions and the correlation between token usage and difficulty and net net, surprise, surprise, like models have got. I think going into next year, that’s going to be really important, especially as you multiply it by the number of steps in an agentic workflow that a model has to take to get to an answer. We are going to care a lot about token efficiency and number of turns efficiency for getting to what we want. Which would you rather have token efficiency or number of turns efficiency? Or like, which is more important to work on?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.