Evidence receipt / recommendation
Published · transcript-backedMentions personal use of Claude.
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case.”
Source trail
Everything needed to verify it.
- Speaker
- Jungwon Byun
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Or Maybe it's the same thing, but I'm thinking especially around this like planning with a world model. Is there something where because you had structured these various lenses and you were able to be more systematic and structured and really agree that this is the framework that we are using, maybe avoiding talking past each other or having some new synthesis, some new insight that you just don't think would have happened in the absence of that? structured approach. This is a partial example. I'm sure Andreas has a better one, but it's actually very timely. I just did this for an important hiring decision where ahead of time we had a rubric designed for how we wanted to, what role we wanted to fill and the role had been through a few different evolutions. We had looked at different personas. There was maybe a disagreement on exactly what type, it's an executive level hire. So some disagreement on exactly what we needed and maybe that changed over time as a company grew during the course of the search. And then at some point with it, over about a month or so ago, I wrote down a framework. And then we did extensive interviews, so many references, lots of back channels. There was just so much information I was getting. And I was starting to develop a take on what we should make or what we should do with this candidate. But I really want to avoid recency bias. And I had a structured project. In this case, I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case. And I created a project where I had all of the different kind of meeting notes across all the different interviews and all the different email threads with this candidate and all the feedback submitted and systematically asked Claude to fill out examples of evidence for every single dimension and it ended up. probably being about 20 different fields, like evidence that this person has consistently hit their goals, evidence that this person can hire a great team, evidence that this person is authentic or is culturally aligned. And I started with evidence and I started piece by piece because I don't trust Claude to fully execute in one go. And there was like a bit of calibration. And then once I had the evidence, which was like hit quota so many years, blah, blah, blah, then was like, okay, what is evidence for against? What decision do I make? How do I rate it on a scale of five? And then all together have this like synthesized point of view. And then I was able to say it to the candidate. I think they really appreciated it as well because they said it was like the greatest kind of comprehensive synthesis of professional validation they had ever received. So that was one case where a compositional structured reasoning with an intentional process and then applied at scale with AI was able to both check my decision-making process and also develop, give someone the gift of something that was like very human and very detailed about them and everything that they had accomplished. My example was going to be much more mundane. I think for me, I've been trying to do this more for just planning my week, where I think about, I have goals and this is what I want to accomplish in the long run this year, this month. And then the question like, I have all these calendar blocks, like this podcast blog, and I need to figure out like there are many different things I could do. What should I do? And which things depend on which other things? It's actually, I think it's a, is a pretty tricky problem to know when you could be spending your time in many different ways, what is worth doing. I've been trying to get to the point where I can use automation as part of my weekly planning and think like more. in a more structured way about for if I want to accomplish my monthly goal, this is where I need to be this week. How much time is that going to take? Is it maybe going to take like 5 hours to write A blog post? When can those five hours happen? And so I think this sort of like backwards chaining. People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.
Named in this claim