High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andreas Stuhlmüller: evaluation

17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

“These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning.”

— Andreas Stuhlmüller

Source trail

Everything needed to verify it.

Speaker
Andreas Stuhlmüller
Attribution
Verified speaker
Claim type
evaluation
Recorded
17 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…I can give a high level take and then I'm sure Andreas will have more a technical response. But one of the projects we worked on in OTT that like many OTT projects was a bit early for its time was we probably built the first language model observability tool. We built this thing called ICE. It stood for Interactive Composition Explorer. And we had this, was, when was this? must have been 2021 or so. And yeah, we anticipated this problem even five years ago that at some point language model traces would get so large, we would not be able to debug them for ourselves. How would we visualize that? How would we maintain oversight? How might we audit that? And so we built this visualization. And one of the things we learned was the best way to troubleshoot is not necessarily to chronologically go through all the steps that the model took. That actually evaluation, so repeating the generation process is one way to check what happened and to build trust in it, but it's inefficient. It doesn't really scale, and it's a very difficult way to check. And so actually what you need often are like a different layer of reasoning checks, almost like logical consistency checks, which are not, let me go through everything you did, steps 1, 2, 3, 4, 5, and see if it was correct. But now let me think about how, for example, like a sensitivity analysis, like how sensitive are my findings to different changes in input parameters, logical consistency checks, things like that. And so that's where I think we need a lot more investment in infrastructures and building kind of independent checks that don't don't rely just on checking chain of thought monitoring. That's my high level take, but yeah, I'm sure you have a more technical version of that answer. Yeah, let me maybe first restate part of what you said in different language. So I think you can either check the process or you can check the outcome. Those are your two options, right? And then when you're checking the outcome, You still want the outcome to somehow contain a certificate that the right reasoning was done. And like that, what can that certificate be, right? It can be like, here's how my conclusion would change if I had a different input. Here's like the things I looked at on the way. Here are like citations to the literature. I do think this is a very underdeveloped field in my mind. So it's In mathematics, it's very developed, right? You can have your formal proof if you want, and that proof is checkable. I don't think it's very developed in more kind of fuzzy domains, and probably partially because we'll just be too. too much work for humans to produce legible certificates, and humans don't even really have great ability to introspect on their own thoughts. But I think in principle, you could, even if you didn't supervise the process, you could produce certificates of reasoning that then let you check the reasoning, even if you didn't check the process that generated that outcome. So I'd be very excited to see more work in that direction. The other clarification I wanted to make is, I think there's For me, it's worth distinguishing like chain of thought and kind of the reasoning process or the chain of process or whatever you want to call it. Because the chain of thought, people often think about what are the thoughts the model writes down, like the reasoning tokens. And then the other, those, how much can you trust them? I don't know. We don't, I think for open AI, we don't even get them anymore these days unless maybe you apply for a special permit to see them. But you do see the tool calls. And I think the tool calls actually are important reasoning facts in and of themselves. Take, for example, let's say I use a model to ask about, like, I download a new paper from archive, the model hasn't seen it before, and ask it some question, ask the model, hey, you know, what are the key results here? And now if the model, I can see which parts of the paper is the model reading, because it, you know, it has a, it's retool that maybe reads, by default the first few lines and then it can scan other parts of the paper. And so sometimes I know the model didn't even look at the methodology section. And that is part of its reasoning that is checkable. I know if it now tells me something about the conclusions of the paper that it really should have relied on the methodology section. I know it didn't do that. That's obviously like a very simple example. These papers are small and so on, but the same thing applies at much larger scale where like you can actually, the process still remains an important waste checkable because we do see the tool calls and the tool calls are an important input into the model's reasoning. Yeah, I like that. How much of this do you think companies are doing today? It's got to be some, but obviously they're not telling us is your on the spectrum from closing in on the sort of thinking that you're doing and the granularity of process supervision that you would like to see on the one end to the purely RLVR with the did you get the right answer in the final answer box or not, just binary signal. How much of this do you think they are doing based on in model behavior?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence