Evidence receipt / preference
Published · transcript-backedAndreas Stuhlmüller: preference
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc.”
Source trail
Everything needed to verify it.
- Speaker
- Andreas Stuhlmüller
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Any particular use cases that you think you with a particular value in that you think other people are maybe sleeping on? I don't know what other people are doing. I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc. You know, when a new day starts, going over like the last day, checking, are there any leftover tasks, like moving them into the right place. So there's a lot of kind of automation that is happening behind, like without me prompting it that probably contributes to those costs being higher. Similarly, for email, I have like pre-elaborates back on like which emails should be auto archived and maybe every hour or so my models check that and go like, okay, you know, let's just archive the emails that Andreas definitely doesn't need to read. And then like on the more kind of user driven side, I do a lot of kind of cross-checking where I run, what are some fun things we could talk about with Nathan and then, okay, but called GPT and Gemini to double-check those things. And I do find the models getting cross-checked by other models often improves the results quite a bit. So that maybe for, I don't know, 1/4 of my use cases that already doubles or triples the cost. So that's another source of additional token spend. Got it. Okay, cool. Other questions I'd love to get your take on is, are we seeing convergence or are we seeing divergence in models? And because one notable feature of elicit today is no model picker, at least from what I've explored recently. So you're making choices and it seems like you clearly think you know best and it would be like, not a good idea, even if people have a favorite model. It would be not a good idea given all the validation and scaffolding that you have to just go in and swap bottle in and out. How do you see this kind of dynamic shaping up? There's again, just such different takes between the model's commodity, scaffolding's all that matters. No, the model's everything. Scaffolding's A complement. They're converging, they're diverging. What is your take on all of that?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.