Evidence receipt / uncertainty
Published · transcript-backedShawn Wang: uncertainty
3 Oct 2024 Latent Space Building AGI in Real Time (OpenAI Dev Day 2024)
“Like, I always want to eval the eval. I don't know if that ever came up. Like, sometimes the evals themselves are wrong, and there's no way for me to tell you.”
Source trail
Everything needed to verify it.
- Speaker
- Shawn Wang
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 3 Oct 2024
- Publisher
- Latent Space
Transcript context
…and we sanitize PII, everything. We have no interest in the actual sensitive data. We just want to have good evaluation on the real use cases. Like, I always want to eval the eval. I don't know if that ever came up. Like, sometimes the evals themselves are wrong, and there's no way for me to tell you. Everyone who is starting with LLM, teaching with LLM, is like, Yeah, evaluation, easy, you know, I've done testing, like, all my life. And then you start to actually be able to eval, understand, like, all the corner cases, And you realize, wow, there's like a whole field in itself. So, yeah, good evaluation is hard and so, yeah. Yeah, yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.