High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Jungwon Byun: belief

11 Apr 2024 Latent Space Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit

“I think we also end up effectively monitoring by trying to evaluate new models as they come out.”

— Jungwon Byun

Source trail

Everything needed to verify it.

Speaker
Jungwon Byun
Attribution
Verified speaker
Claim type
belief
Recorded
11 Apr 2024
Publisher
Latent Space

Transcript context

…And then things like hallucination rate where I think there, the really important thing is training time. So we care a lot about having our own internal benchmarks for model development that reflect the distribution of user queries so that we can know ahead of time how well is the model going to perform on different types of tasks. So the tasks being summarization, question answering, given a paper, ranking. And for each of those, we want to know what's the distribution of things the model is going to see so that we can have well-calibrated predictions on how well the model is going to do in production. And I think, yeah, there's some chance that there's distribution shift and actually the things users enter are going to be different. But I think that's much less important than getting the kind of training right and having very high quality, well-vetted data sets at training time. I think we also end up effectively monitoring by trying to evaluate new models as they come out. And so that kind of prompts us to go through our eval suite every couple of months. And every time a new model comes out, we have to see how is this performing relative to production and what we currently have. Yeah. I mean, since we're on this topic, any new models that have really caught your eye this year?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence