Evidence receipt / uncertainty
Published · transcript-backedShawn Wang: uncertainty
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“We don’t know for a fact that it’s like eight runs and then with the LM judge on top.”
Source trail
Everything needed to verify it.
- Speaker
- Shawn Wang
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…Oh yeah. Cause GPT pros are rumored. We don’t know for a fact that it’s like eight runs and then with the LM judge on top. Yeah. So we saw a big jump in, this is accuracy. So this is just percent that they get, uh, correct and Gemini three pro knew a lot more than the other models. And so big jump in accuracy. But relatively no change between the Google Gemini models, between releases. And the hallucination rate. Exactly. And so it’s likely due to just kind of different post-training recipe, between the, the Claude models. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.