Evidence receipt / belief
Published · transcript-backedJungwon Byun: belief
11 Apr 2024 Latent Space Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
“I think we also end up effectively monitoring by trying to evaluate new models as they come out.”
Source trail
Everything needed to verify it.
- Speaker
- Jungwon Byun
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 11 Apr 2024
- Publisher
- Latent Space
Transcript context
…And then things like hallucination rate where I think there, the really important thing is training time. So we care a lot about having our own internal benchmarks for model development that reflect the distribution of user queries so that we can know ahead of time how well is the model going to perform on different types of tasks. So the tasks being summarization, question answering, given a paper, ranking. And for each of those, we want to know what's the distribution of things the model is going to see so that we can have well-calibrated predictions on how well the model is going to do in production. And I think, yeah, there's some chance that there's distribution shift and actually the things users enter are going to be different. But I think that's much less important than getting the kind of training right and having very high quality, well-vetted data sets at training time. I think we also end up effectively monitoring by trying to evaluate new models as they come out. And so that kind of prompts us to go through our eval suite every couple of months. And every time a new model comes out, we have to see how is this performing relative to production and what we currently have. Yeah. I mean, since we're on this topic, any new models that have really caught your eye this year?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.