High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andreas Stuhlmüller: evaluation

17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

“Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process.”

— Andreas Stuhlmüller

Source trail

Everything needed to verify it.

Speaker
Andreas Stuhlmüller
Attribution
Verified speaker
Claim type
evaluation
Recorded
17 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…Could you give me a double click into what exactly you had in mind when you were talking about how you were going to improve the quality of reasoning, how that compares to the more and more RL on top of a chain of thought paradigm that everybody now is very familiar with, and then how that has maybe enabled what you're doing with Elicit, maybe in some ways competed with it, and how it has shaped your business and product strategy to go in this more principled, a little less bitter pilled direction with what you're trying to build. I like bitter pilled. Yeah, so last time we talked, we talked a lot about process supervision, right? How do you know that a system is doing good research for you? can either look at the output and be like, that looks good to me, or you can look at the process it went through, you know, what papers did it look at, why did it look at them, how did it choose what to do? And Back in the day, when we were young and naive, you all thought, it would be better to know that things are correct for the right reasons, that the process was good. And so we leaned a lot into that. And in many ways, I think the justification for that has still barred out. So let me talk about a quick anecdotal experiment we ran, I think, just two days ago, where we told a few research agents, hey, look through and analyze about 100 papers on toxicology risk for a particular type of cancer drug. And then, so we tell that to Claude, we tell that to ChatGPT, tell that to Elicit, and then we ask, hey, how many papers did you actually analyze? And then as the models like to do, they're like, you know, that's a fair and important question to ask. Let me be direct. I did not analyze 100 papers. You're right to push back. I didn't do it. And that I think that is like a failure of process in many ways. Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process. The models are trained to produce outputs that where if you didn't check and you just looked at the result that said, hey, here's my analysis, you wouldn't have caught that. I think the fundamental problem still exists. And then the question is like, what do we do about it? And to what extent does the solution look like better checking of the outputs versus checking of the process in various ways. I actually think this question is still open. I can speak for what Elicit does and maybe briefly for what the rest of the ecosystem does. Elicit addresses this by, we have like a little kind of domain-specific language that the research agent can write that orchestrates other calls to agents. Screen all these papers, then extract data from all these papers, and it runs the, you run the process, you know the process does what it said it was going to do. I think actually the model companies are going a little bit in that direction too. I think I haven't been following it extremely closely, but I think Anthropic recently launched like a workflows feature that has similar properties. And so even though I think on some level, yeah, all the models are outcome trained and that's why we see these like quite ridiculous artifacts of models being like, oh, sorry, I didn't do it. I think the problem still exists and like one level up from that, people are trying to patch it. that's where I see ourselves as being at in this game. Yeah, I think it's worth emphasizing that a little bit more because it's a core part of how Elicit is built differently and therefore what you can use Elicit for. A lot of, like you said, it's been 2 years since we chatted. A lot has happened in that time. A big part of what we've been working on for the last year is just rebuilding Elicit on top of this much more agentic infrastructure. We started working on that in about March or April of 2025. In retrospect, it was maybe a little bit early, but at the time it felt quite late because obviously a lot of advances happened over the last two, three quarters. But a big part of what we spent the last year thinking about was designing how do we preserve these benefits of transparency, systematicity at scale without losing the flexibility and raw power of these models? I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature. But you don't want to be overly deterministic because then you run into the bitter lesson issue, right? So like threading the needle is what we're always struggling with. So a lot of our time Last year was spent on this kind of technical design question, and we decided to design our own programming language to solve this problem so that the models could run reasoning computation, be able to call these reasoning primitives at scale in a more trustworthy way. And what we were trying to accomplish for our end users was the ability to say, you can run this process with the model over 10,000 objects, 10,000 documents, 10,000 drugs, 10,000 targets, genes, whatever, and the same process will be applied to #5 as number 9,999. And there's just no other model that can make that guarantee. And unfortunately, there are lots of models that claim that they have done that or can do that and are just completely wrong about it. So when we think about who are our users, how do we build a differentiated product that really meets their needs, what use cases can we enable? We are powering people who want to be able to rigorously synthesize evidence at a very large scale and get every single thing right all the way to the nth degree. That's a very different interaction than just riffing with a model. And so that's, and we support some of these lightweight use cases too, but I think there it's almost more like getting to parity. And I think where Elicit's really differentiated is that trust at scale.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence