High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Shawn Wang: uncertainty

8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

“All these nuances, like, once you dig into benchmarks, you’re like, I don’t know how anyone believes the numbers on all these things.”

— Shawn Wang

Source trail

Everything needed to verify it.

Speaker
Shawn Wang
Attribution
Verified speaker
Claim type
uncertainty
Recorded
8 Jan 2026
Publisher
Latent Space

Transcript context

…It depends what you’re looking at, right? Because you can, if you’re trying to see whether or not it can solve a particular type of reasoning problem, and you don’t want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how unanswered. But these days, it’s mostly less of a problem. Like, if you instruct a model and give it examples of what the answers should look like, it can get the answers in your format, and then you can do, like, a simple regex. Yeah, yeah. And then there’s other questions around, I guess, sometimes if you have a multiple choice question, sometimes there’s a bias towards the first answer, so you have to randomize the responses. All these nuances, like, once you dig into benchmarks, you’re like, I don’t know how anyone believes the numbers on all these things. It’s so dark magic. You’ve also got, like… You’ve got, like, the different degrees of variance in different benchmarks, right? Yeah. So, if you run four-question multi-choice on a modern reasoning model at the temperatures suggested by the labs for their own models, the variance that you can see on a four-question multi-choice eval is pretty enormous if you only do a single run of it and it has a small number of questions, especially. So, like, one of the things that we do is run an enormous number of all of our evals when we’re developing new ones and doing upgrades to our intelligence index to bring in new things. Yeah. So, that we can dial in the right number of repeats so that we can get to the 95% confidence intervals that we’re comfortable with so that when we pull that together, we can be confident in intelligence index to at least as tight as, like, a plus or minus one at a 95% confidence. Yeah.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence