High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

The Cognitive Revolution / episode intelligence

Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

17 Jun 2026 22 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 10recommendation 5evaluation 5disagreement 1preference 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

22 published records

02 / recommendation

Mentions personal use of systematic literature view API.

“When I run systematic literature views, often I kind of use the systematic literature view API and I iterate using like various other models on the protocol for a while, and then I run it in the background and retrieve it.”
Publisher
The Cognitive Revolution

04 / belief

People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.

“People do it like informally, but I think there is a lot of kind of constraint satisfaction and like propagation of constraints that it's pretty tricky for humans and that I think the models will help us with that.”
Publisher
The Cognitive Revolution

05 / belief

For the time being, we still have some agency over this process, and I think that's a great reminder to maintain an ownership mindset and hold ourselves accountable to doing our very best work and not getting lazy and letting the AIs lead us around.

“For the time being, we still have some agency over this process, and I think that's a great reminder to maintain an ownership mindset and hold ourselves accountable to doing our very best work and not getting lazy and letting the AIs lead us around.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

06 / belief

You could the decomposition process like multiple times and check for consistency, which I think you're suggesting something like that might be going on.

“You could the decomposition process like multiple times and check for consistency, which I think you're suggesting something like that might be going on.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution

07 / belief

I feel like there's always an explore-exploit trade-off, and you want to navigate those two thoughtfully, and sometimes you want to do one, sometimes you want to do the other. And maybe the cost of exploring has gone down a little bit, but I think there are still some things where you know you have to get it right the first time, or it's still, it's not all software engineering has literally gone to zero.

“I feel like there's always an explore-exploit trade-off, and you want to navigate those two thoughtfully, and sometimes you want to do one, sometimes you want to do the other. And maybe the cost of exploring has gone down a little bit, but I think there are still some things where you know you have to get it right the first time, or it's still, it's not all software engineering has literally gone to zero.”
Speaker
Jungwon Byun
Publisher
The Cognitive Revolution

08 / belief

I still find, especially as a company that is trying to be extremely mission focused and how do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality, I think by default, you're probably just not going to accomplish that.

“I still find, especially as a company that is trying to be extremely mission focused and how do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality, I think by default, you're probably just not going to accomplish that.”
Publisher
The Cognitive Revolution

12 / belief

The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data.

“The reason is I think a lot of tasks still involve kind of multiple models orchestrated in a way that we think makes the most sense, having like particular models that are good at screening papers or extracting data.”
Publisher
The Cognitive Revolution

13 / disagreement

I think sometimes people like come to us and are like, well, who's going to win in AI for science? And I think that's just an absurd thing to say because science is such a big space.

“I think sometimes people like come to us and are like, well, who's going to win in AI for science? And I think that's just an absurd thing to say because science is such a big space.”
Publisher
The Cognitive Revolution

14 / preference

I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc.

“I have found, as mentioned earlier, I think I do have like a lot of use in planning, keeping my calendar in sync with my personal journaling system, in sync with my to-dos, making sure everything is coherent with my longer range planning doc.”
Publisher
The Cognitive Revolution

15 / recommendation

But now let me think about how, for example, like a sensitivity analysis, like how sensitive are my findings to different changes in input parameters, logical consistency checks, things like that. And so that's where I think we need a lot more investment in infrastructures and building kind of independent checks that don't don't rely just on checking chain of thought monitoring.

“But now let me think about how, for example, like a sensitivity analysis, like how sensitive are my findings to different changes in input parameters, logical consistency checks, things like that. And so that's where I think we need a lot more investment in infrastructures and building kind of independent checks that don't don't rely just on checking chain of thought monitoring.”
Speaker
Jungwon Byun
Publisher
The Cognitive Revolution

16 / evaluation

I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time.

“I think with software engineering, if it's the case that 80% of the time when the model says this is like an automatically reviewable feature than it is actually is, then that's not good enough because we don't want to break production 20% of the time.”
Publisher
The Cognitive Revolution

17 / evaluation

I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal.

“I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal.”
Publisher
The Cognitive Revolution

18 / evaluation

I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature.

“I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature.”
Speaker
Jungwon Byun
Publisher
The Cognitive Revolution

19 / evaluation

These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning.

“These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning.”
Publisher
The Cognitive Revolution

20 / evaluation

Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process.

“Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process.”
Publisher
The Cognitive Revolution

21 / recommendation

Mentions personal use of Claude Max. Mentions personal use of Codex Pro.

“And certainly if you were to triple it from there, you'd be getting into something on the order of magnitude of parity with human headcount. What are you doing with it all too? Because I use my $200 Claude Max and my Codex Pro and I honestly don't even hit my limits that often.”
Speaker
Nathan Labenz
Publisher
The Cognitive Revolution
Search evidence