Evidence receipt / evaluation
Published · transcript-backedJungwon Byun: evaluation
17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
“I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature.”
Source trail
Everything needed to verify it.
- Speaker
- Jungwon Byun
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 17 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…I like bitter pilled. Yeah, so last time we talked, we talked a lot about process supervision, right? How do you know that a system is doing good research for you? can either look at the output and be like, that looks good to me, or you can look at the process it went through, you know, what papers did it look at, why did it look at them, how did it choose what to do? And Back in the day, when we were young and naive, you all thought, it would be better to know that things are correct for the right reasons, that the process was good. And so we leaned a lot into that. And in many ways, I think the justification for that has still barred out. So let me talk about a quick anecdotal experiment we ran, I think, just two days ago, where we told a few research agents, hey, look through and analyze about 100 papers on toxicology risk for a particular type of cancer drug. And then, so we tell that to Claude, we tell that to ChatGPT, tell that to Elicit, and then we ask, hey, how many papers did you actually analyze? And then as the models like to do, they're like, you know, that's a fair and important question to ask. Let me be direct. I did not analyze 100 papers. You're right to push back. I didn't do it. And that I think that is like a failure of process in many ways. Because you're like, I told you what to do, you didn't do it. And so I think the reason for that is because the models are not trained on process. The models are trained to produce outputs that where if you didn't check and you just looked at the result that said, hey, here's my analysis, you wouldn't have caught that. I think the fundamental problem still exists. And then the question is like, what do we do about it? And to what extent does the solution look like better checking of the outputs versus checking of the process in various ways. I actually think this question is still open. I can speak for what Elicit does and maybe briefly for what the rest of the ecosystem does. Elicit addresses this by, we have like a little kind of domain-specific language that the research agent can write that orchestrates other calls to agents. Screen all these papers, then extract data from all these papers, and it runs the, you run the process, you know the process does what it said it was going to do. I think actually the model companies are going a little bit in that direction too. I think I haven't been following it extremely closely, but I think Anthropic recently launched like a workflows feature that has similar properties. And so even though I think on some level, yeah, all the models are outcome trained and that's why we see these like quite ridiculous artifacts of models being like, oh, sorry, I didn't do it. I think the problem still exists and like one level up from that, people are trying to patch it. that's where I see ourselves as being at in this game. Yeah, I think it's worth emphasizing that a little bit more because it's a core part of how Elicit is built differently and therefore what you can use Elicit for. A lot of, like you said, it's been 2 years since we chatted. A lot has happened in that time. A big part of what we've been working on for the last year is just rebuilding Elicit on top of this much more agentic infrastructure. We started working on that in about March or April of 2025. In retrospect, it was maybe a little bit early, but at the time it felt quite late because obviously a lot of advances happened over the last two, three quarters. But a big part of what we spent the last year thinking about was designing how do we preserve these benefits of transparency, systematicity at scale without losing the flexibility and raw power of these models? I think that's the core design question we've always wrestled with because when you want to deploy these models at scale for really high stakes decisions, you need to be able, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature. But you don't want to be overly deterministic because then you run into the bitter lesson issue, right? So like threading the needle is what we're always struggling with. So a lot of our time Last year was spent on this kind of technical design question, and we decided to design our own programming language to solve this problem so that the models could run reasoning computation, be able to call these reasoning primitives at scale in a more trustworthy way. And what we were trying to accomplish for our end users was the ability to say, you can run this process with the model over 10,000 objects, 10,000 documents, 10,000 drugs, 10,000 targets, genes, whatever, and the same process will be applied to #5 as number 9,999. And there's just no other model that can make that guarantee. And unfortunately, there are lots of models that claim that they have done that or can do that and are just completely wrong about it. So when we think about who are our users, how do we build a differentiated product that really meets their needs, what use cases can we enable? We are powering people who want to be able to rigorously synthesize evidence at a very large scale and get every single thing right all the way to the nth degree. That's a very different interaction than just riffing with a model. And so that's, and we support some of these lightweight use cases too, but I think there it's almost more like getting to parity. And I think where Elicit's really differentiated is that trust at scale. Yeah, I should always remind myself and be clear that none of these things are really true binaries in the sense that it would be wrong to characterize Elicit as being not bitter pilled because I remember last time one of your big principles was how can we allow you to spend more money to buy more compute to get better results? And that's almost a restatement of the bitter lesson in some way. And at the same time, the frontier companies, as you said, are doing some of this. And I assume that in their research agents, in particular, they're deep research products. They are presumably doing some sort of at least rubric-based reward on the final reports that the models are outputting, albeit, if I understand correctly, still at least intending to avoid putting any optimization pressure on the chain of thought. So there's always little aspects of gray to this. How would you say the So who are the customers now? There are these deep research things. I use those pretty frequently. Who is the sweet spot that is like, a deep research agent isn't enough for me. I really want to go way bigger, way more systematic, be very sure that I'm performing the same analysis in a way that I can count on. Who are those customers now that you're finding product market fit with?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.