High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Andreas Stuhlmüller: belief

11 Apr 2024 Latent Space Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit

“I think in some cases, we also use two different models for the uncertainty estimates than for the question answering.”

— Andreas Stuhlmüller

Source trail

Everything needed to verify it.

Speaker
Andreas Stuhlmüller
Attribution
Verified speaker
Claim type
belief
Recorded
11 Apr 2024
Publisher
Latent Space

Transcript context

…We found it to be pretty calibrated. It varies on the model. I think in some cases, we also use two different models for the uncertainty estimates than for the question answering. So one model would say, here's my chain of thought, here's my answer. And then a different type of model. Let's say the first model is Llama, and let's say the second model is GPT-3.5. And then the second model just looks over the results and is like, okay, how confident are you in this? And I think sometimes using a different model can be better than using the same model. Yeah. On the topic of models, evaluating models, obviously you can do that all day long. What's your budget? Because your queries fan out a lot. And then you have models evaluating models. One person typing in a question can lead to a thousand calls.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence