Evidence receipt / belief
Published · transcript-backedAndreas Stuhlmüller: belief
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.”
Source trail
Everything needed to verify it.
- Speaker
- Andreas Stuhlmüller
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…This is a partial example. I'm sure Andreas has a better one, but it's actually very timely. I just did this for an important hiring decision where ahead of time we had a rubric designed for how we wanted to, what role we wanted to fill and the role had been through a few different evolutions. We had looked at different personas. There was maybe a disagreement on exactly what type, it's an executive level hire. So some disagreement on exactly what we needed and maybe that changed over time as a company grew during the course of the search. And then at some point with it, over about a month or so ago, I wrote down a framework. And then we did extensive interviews, so many references, lots of back channels. There was just so much information I was getting. And I was starting to develop a take on what we should make or what we should do with this candidate. But I really want to avoid recency bias. And I had a structured project. In this case, I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case. And I created a project where I had all of the different kind of meeting notes across all the different interviews and all the different email threads with this candidate and all the feedback submitted and systematically asked Claude to fill out examples of evidence for every single dimension and it ended up. probably being about 20 different fields, like evidence that this person has consistently hit their goals, evidence that this person can hire a great team, evidence that this person is authentic or is culturally aligned. And I started with evidence and I started piece by piece because I don't trust Claude to fully execute in one go. And there was like a bit of calibration. And then once I had the evidence, which was like hit quota so many years, blah, blah, blah, then was like, okay, what is evidence for against? What decision do I make? How do I rate it on a scale of five? And then all together have this like synthesized point of view. And then I was able to say it to the candidate. I think they really appreciated it as well because they said it was like the greatest kind of comprehensive synthesis of professional validation they had ever received. So that was one case where a compositional structured reasoning with an intentional process and then applied at scale with AI was able to both check my decision-making process and also develop, give someone the gift of something that was like very human and very detailed about them and everything that they had accomplished. My example was going to be much more mundane. I think for me, I've been trying to do this more for just planning my week, where I think about, I have goals and this is what I want to accomplish in the long run this year, this month. And then the question like, I have all these calendar blocks, like this podcast blog, and I need to figure out like there are many different things I could do. What should I do? And which things depend on which other things? It's actually, I think it's a, is a pretty tricky problem to know when you could be spending your time in many different ways, what is worth doing. I've been trying to get to the point where I can use automation as part of my weekly planning and think like more. in a more structured way about for if I want to accomplish my monthly goal, this is where I need to be this week. How much time is that going to take? Is it maybe going to take like 5 hours to write A blog post? When can those five hours happen? And so I think this sort of like backwards chaining. People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that. Yeah, that's cool. I think of that as building your own harness in a way, which is something I'm thinking about for myself too. Like how can I build up structures around me to keep steering me in the right direction, feeding me the information I need, and hopefully helping me become my best self, use my time as well as I possibly can by setting me up for success as much as AIs can do that. Do you find that you are following it? Do you find your, like, how good is it? And are you actually living by it yet? Or is it still, maybe I'll, maybe next week when it gets a little better, I'll actually do the plan that it gives me.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.