Evidence receipt / belief
Published · transcript-backedAndreas Stuhlmüller: belief
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data.”
Source trail
Everything needed to verify it.
- Speaker
- Andreas Stuhlmüller
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Got it. Okay, cool. Other questions I'd love to get your take on is, are we seeing convergence or are we seeing divergence in models? And because one notable feature of elicit today is no model picker, at least from what I've explored recently. So you're making choices and it seems like you clearly think you know best and it would be like, not a good idea, even if people have a favorite model. It would be not a good idea given all the validation and scaffolding that you have to just go in and swap bottle in and out. How do you see this kind of dynamic shaping up? There's again, just such different takes between the model's commodity, scaffolding's all that matters. No, the model's everything. Scaffolding's A complement. They're converging, they're diverging. What is your take on all of that? Yeah, I keep being shocked by how much the models are converging. I think it's really, I guess I should stop being shocked at this point because I'm just not updating, but it is a really interesting and surprising fact about the world that the models are so similar. I do, that's not the reason why we don't offer a model picker. The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data. And I think the differences, even though the models are so similar, I think the differences are important in subtle ways. So for example, I think like people hate to hate on Gemini and so do I, but when we evaluated like Claude, Opus, I think this was like 4.5 against like Gemini 3 Pro at the time. I think Opus like did like better on extraction accuracy, but if you check that, you know what fraction of claims are directly supported by the evidence. I think Gemini actually beat it by like at least 5% or so. there's, I think it's still, yeah, I think the models are still like, I don't know, like micro jagged enough that you can't say, oh yeah, this model is clearly the best. You should just use that. And so as a user, I don't want to put that on our users for the most part. I think mostly what our users pay for is for us to do the work of figuring out what models are good at what kind of thing and making sure those models are actually getting used in those places. Yeah, interesting. So there is a place for Gemini in Elicit today.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.