High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Shawn Wang: belief

8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

“I think there’s a general field of calibration as well, like the confidence in your answer versus the rightness of the answer.”

— Shawn Wang

Source trail

Everything needed to verify it.

Speaker
Shawn Wang
Attribution
Verified speaker
Claim type
belief
Recorded
8 Jan 2026
Publisher
Latent Space

Transcript context

…Yeah, that’s fair. There’s a lot of fun stuff to dive into. Maybe so we can hit past all the, like, we have lots and lots of emails and stuff. The interesting ones to talk about today that would be great to bring up are a few of our recent things, I think, that probably not many people will be familiar with yet. So first one of those is our omniscience index. So this one is a little bit different to most of the intelligence evils that we’ve run. We built it specifically to look at the embedded knowledge in the models and to test hallucination by looking at when the model doesn’t know the answer, so not able to get it correct, what’s its probability of saying, I don’t know, or giving an incorrect answer. So the metric that we use for omniscience goes from negative 100 to positive 100. Because we’re simply taking off a point if you give an incorrect answer to the question. We’re pretty convinced that this is an example of where it makes most sense to do that, because it’s strictly more helpful to say, I don’t know, instead of giving a wrong answer to factual knowledge question. And one of our goals is to shift the incentive that evils create for models and the labs creating them to get higher scores. And almost every evil across all of AI up until this point, it’s been graded by simple percentage correct as the main metric, the main thing that gets hyped. And so you should take a shot at everything. There’s no incentive to say, I don’t know. So we did that for this one here. I think there’s a general field of calibration as well, like the confidence in your answer versus the rightness of the answer. Yeah, we completely agree. Yeah. Yeah. On that. And one reason that we didn’t do that is because. Or put that into this index is that we think that the, the way to do that is not to ask the models how confident they are.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence