Evidence receipt / uncertainty
Published · transcript-backedShawn Wang: uncertainty
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“All these nuances, like, once you dig into benchmarks, you’re like, I don’t know how anyone believes the numbers on all these things.”
Source trail
Everything needed to verify it.
- Speaker
- Shawn Wang
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…It depends what you’re looking at, right? Because you can, if you’re trying to see whether or not it can solve a particular type of reasoning problem, and you don’t want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how unanswered. But these days, it’s mostly less of a problem. Like, if you instruct a model and give it examples of what the answers should look like, it can get the answers in your format, and then you can do, like, a simple regex. Yeah, yeah. And then there’s other questions around, I guess, sometimes if you have a multiple choice question, sometimes there’s a bias towards the first answer, so you have to randomize the responses. All these nuances, like, once you dig into benchmarks, you’re like, I don’t know how anyone believes the numbers on all these things. It’s so dark magic. You’ve also got, like… You’ve got, like, the different degrees of variance in different benchmarks, right? Yeah. So, if you run four-question multi-choice on a modern reasoning model at the temperatures suggested by the labs for their own models, the variance that you can see on a four-question multi-choice eval is pretty enormous if you only do a single run of it and it has a small number of questions, especially. So, like, one of the things that we do is run an enormous number of all of our evals when we’re developing new ones and doing upgrades to our intelligence index to bring in new things. Yeah. So, that we can dial in the right number of repeats so that we can get to the 95% confidence intervals that we’re comfortable with so that when we pull that together, we can be confident in intelligence index to at least as tight as, like, a plus or minus one at a 95% confidence. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.