← All source episodes The Cognitive Revolution / episode intelligence
Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
17 Jun 2026 22 published claims 3 attributable people
Speakers in the public record
Claim mix
belief 10recommendation 5evaluation 5disagreement 1preference 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
22 published records
“What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
- Publisher
- The Cognitive Revolution
“When I run systematic literature views, often I kind of use the systematic literature view API and I iterate using like various other models on the protocol for a while, and then I run it in the background and retrieve it.”
- Publisher
- The Cognitive Revolution
“I think knowing when the models are making you better or worse at decision making is actually pretty subtle.”
- Publisher
- The Cognitive Revolution
“People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.”
- Publisher
- The Cognitive Revolution
“For the time being, we still have some agency over this process, and I think that's a great reminder to maintain an ownership mindset and hold ourselves accountable to doing our very best work and not getting lazy and letting the AIs lead us around.”
- Publisher
- The Cognitive Revolution
“You could the decomposition process like multiple times and check for consistency, which I think you're suggesting something like that might be going on.”
- Publisher
- The Cognitive Revolution
“I feel like there's always an explore-exploit trade-off, and you want to navigate those two thoughtfully, and sometimes you want to do one, sometimes you want to do the other. And maybe the cost of exploring has gone down a little bit, but I think there are still some things where you know you have to get it right the first time, or it's still, it's not all software engineering has literally gone to zero.”
- Publisher
- The Cognitive Revolution
“I still find, especially as a company that is trying to be extremely mission focused and how do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality, I think by default, you're probably just not going to accomplish that.”
- Publisher
- The Cognitive Revolution
“I think of that as building your own harness in a way, which is something I'm thinking about for myself too.”
- Publisher
- The Cognitive Revolution
“I think even though I'm fairly on top of what's going on, our EVOS team is even more on top of it.”
- Publisher
- The Cognitive Revolution
“I think step one is just to use the representations people already find useful. And SQL databases are a pretty useful thing.”
- Publisher
- The Cognitive Revolution
“The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data.”
- Publisher
- The Cognitive Revolution
“I think sometimes people like come to us and are like, well, who's going to win in AI for science? And I think that's just an absurd thing to say because science is such a big space.”
- Publisher
- The Cognitive Revolution
“I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc.”
- Publisher
- The Cognitive Revolution
“But now let me think about how, for example, like a sensitivity analysis, like how sensitive are my findings to different changes in input parameters, logical consistency checks, things like that. And so that's where I think we need a lot more investment in infrastructures and building kind of independent checks that don't don't rely just on checking chain of thought monitoring.”
- Publisher
- The Cognitive Revolution
“I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time.”
- Publisher
- The Cognitive Revolution
“I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal.”
- Publisher
- The Cognitive Revolution
“I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature.”
- Publisher
- The Cognitive Revolution
“These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning.”
- Publisher
- The Cognitive Revolution
“Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process.”
- Publisher
- The Cognitive Revolution
“And certainly if you were to triple it from there, you'd be getting into something on the order of magnitude of parity with human headcount. What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
- Publisher
- The Cognitive Revolution
“I use Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, or I don't use it for that case.”
- Publisher
- The Cognitive Revolution