High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Shawn Wang: uncertainty

8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

“We don’t know for a fact that it’s like eight runs and then with the LM judge on top.”

— Shawn Wang

Source trail

Everything needed to verify it.

Speaker
Shawn Wang
Attribution
Verified speaker
Claim type
uncertainty
Recorded
8 Jan 2026
Publisher
Latent Space

Transcript context

…Oh yeah. Cause GPT pros are rumored. We don’t know for a fact that it’s like eight runs and then with the LM judge on top. Yeah. So we saw a big jump in, this is accuracy. So this is just percent that they get, uh, correct and Gemini three pro knew a lot more than the other models. And so big jump in accuracy. But relatively no change between the Google Gemini models, between releases. And the hallucination rate. Exactly. And so it’s likely due to just kind of different post-training recipe, between the, the Claude models. Yeah.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence