High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andreas Stuhlmüller: evaluation

17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

“I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal.”

— Andreas Stuhlmüller

Source trail

Everything needed to verify it.

Speaker
Andreas Stuhlmüller
Attribution
Verified speaker
Claim type
evaluation
Recorded
17 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, that reminds me of something that you recently wrote around Elicit's focus being on becoming the best at reducing these big kind of fuzzy, hard to verify tasks to sets or graphs perhaps of easy to verify tasks. And I'd love to hear a little bit more, you've alluded to it there, but I'd love to hear a little bit more about like how you're going about that and also like how complete you think that process can be. I guess everywhere I look these days, I feel like I see this same shape of a really interesting question that I don't know what to make of, which is like basically how do we get high level guarantees or conclusions or insights from low level steps, right? And this is like in biology, I might be able to say I've got all these proteins or these genes are being expressed at this level, but I know what's going to happen next at the cell or the tissue or the organism level. No, right? I don't. And similarly, it seems with these kind of big judgment calls, should I prioritize this drug or that drug, we can break it down and get systematic. But it's not clear to me if there's, if it kind of, But how close it gets to something where I'm like, yes, okay, I can really buy in and trust that versus is there something that kind of sits above all those steps still that is, it emergent or is it just somehow lost or people are doing some sort of meta cognitive work that's hard to capture that is maybe still like very critical to actually being effective at these tasks. So I guess, again, tell me everything, but I'm really struggling a lot with this like, Formal methods, formal verifications is another area where I see this, where we can make all these low-level guarantees about this error, that error can't happen, but is the system itself going to behave how we wanted it to? I still don't know in a lot of cases how we make that leap. So very interested in your take on these sort of, I usually think of it as like laddering up low-level things to high-level conclusions where you're approaching it from the other direction, which is interesting unto itself. Yeah, again, tell me everything. maybe I guess to quickly restate, what is the, why do we want to reduce hard to verify tasks to easy to verify tasks? It's because AI currently can be trained on easy to verify tasks. We know it's like extremely good at RLVR coding. It's a tough math tasks like this. And it's quite, I would say like quite weak at a lot of fuzzy tasks. So when I, notice it all the time when I try to use The models to help me in plan our company strategy, for example, I think they're surprisingly useless. So even though they have access to all my contacts, they're really quite good at being like, you know, let me pull in the data, let me pull in your email, your Slack. I still find that I can't, they don't get it. Or like in an important way, this is related to what we said earlier about like they're too easy to push around. It doesn't feel like they're building up a coherent model of what's going on. And I think an important The reason for that is, that is a hard to check tasks. Then, okay, what to do? I think it depends a little bit on like what your situation is, whether you're trying to create a reward signal that you can train the models on, or whether you're trying to kind of do verification and checking for the purpose of understanding whether an already trained model can be trusted in a situation or how to refine its behavior. I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal. And so it's not enough to do spot checks and be like, hey, here are some cases where we can verify that, for example, your company strategy was, I don't know, incompatible with some claim you made earlier. Whereas if the goal is to take an already trained model and get it and understand how good is it exactly, or did it make some fairly obvious mistakes? Can I find places where it can improve? I think then your reward signal doesn't need to be bulletproof or your way of getting some easy to verify aspects of the hard to verify task. You can make more kind of stepwise progress, I would think. And so I think our situation, like we are not currently trying to train like a foundation model from scratch or even like post-train a model on this particular aspect. And so our situation is more like, how do we get to the point where we can check many important properties of tasks? Like are the claims the model makes internally consistent? Is it the case that, you know, if we broke it down in different ways, it will end up at the same conclusions and so on? I think that is like a fairly tractable project, I think. And the project of, how do we fully reduce a high level task like company strategy into individual components that are all like formally verifiable, I think is a much rougher prospect. Not isn't to say it's impossible, but it's like less of a, you have less incremental kind of feedback signal that you're on the right track there, I would say. So Tell me a little bit more about what you're doing in practice. You said you're not trying to post train a model. I know in the past there was a decent amount of fine tuning though, at least for kind of specific tasks. I'm curious if there's still a fine tuning element to it. And then there's a lot of different ways you could think about spending a lot of tokens to try to get at this. You could the decomposition process like multiple times and check for consistency, which I think you're suggesting something like that might be going on. You could do a more iterative thing where you get the AI to give you an output and then have some kind of specialist prompts or perhaps even specialist fine tune models come in and assess in various ways and then give it feedback and then let it reason some more and try to improve on what it just did. And we do see some of that stuff. I just talked to some OpenAI forward deployed engineers who are basically using that process to climb the hill on filing accuracy and that seems to be going quite well for them. So what techniques are you finding to be in practice most effective today?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence