High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

17 Jun 2026 The Cognitive Revolution Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

“You could the decomposition process like multiple times and check for consistency, which I think you're suggesting something like that might be going on.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
17 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…maybe I guess to quickly restate, what is the, why do we want to reduce hard to verify tasks to easy to verify tasks? It's because AI currently can be trained on easy to verify tasks. We know it's like extremely good at RLVR coding. It's a tough math tasks like this. And it's quite, I would say like quite weak at a lot of fuzzy tasks. So when I, notice it all the time when I try to use The models to help me in plan our company strategy, for example, I think they're surprisingly useless. So even though they have access to all my contacts, they're really quite good at being like, you know, let me pull in the data, let me pull in your email, your Slack. I still find that I can't, they don't get it. Or like in an important way, this is related to what we said earlier about like they're too easy to push around. It doesn't feel like they're building up a coherent model of what's going on. And I think an important The reason for that is, that is a hard to check tasks. Then, okay, what to do? I think it depends a little bit on like what your situation is, whether you're trying to create a reward signal that you can train the models on, or whether you're trying to kind of do verification and checking for the purpose of understanding whether an already trained model can be trusted in a situation or how to refine its behavior. I think if you're trying to create a reward signal, that's pretty rough because The models are going to optimize pretty hard against your signal. And so it's not enough to do spot checks and be like, hey, here are some cases where we can verify that, for example, your company strategy was, I don't know, incompatible with some claim you made earlier. Whereas if the goal is to take an already trained model and get it and understand how good is it exactly, or did it make some fairly obvious mistakes? Can I find places where it can improve? I think then your reward signal doesn't need to be bulletproof or your way of getting some easy to verify aspects of the hard to verify task. You can make more kind of stepwise progress, I would think. And so I think our situation, like we are not currently trying to train like a foundation model from scratch or even like post-train a model on this particular aspect. And so our situation is more like, how do we get to the point where we can check many important properties of tasks? Like are the claims the model makes internally consistent? Is it the case that, you know, if we broke it down in different ways, it will end up at the same conclusions and so on? I think that is like a fairly tractable project, I think. And the project of, how do we fully reduce a high level task like company strategy into individual components that are all like formally verifiable, I think is a much rougher prospect. Not isn't to say it's impossible, but it's like less of a, you have less incremental kind of feedback signal that you're on the right track there, I would say. So Tell me a little bit more about what you're doing in practice. You said you're not trying to post train a model. I know in the past there was a decent amount of fine tuning though, at least for kind of specific tasks. I'm curious if there's still a fine tuning element to it. And then there's a lot of different ways you could think about spending a lot of tokens to try to get at this. You could the decomposition process like multiple times and check for consistency, which I think you're suggesting something like that might be going on. You could do a more iterative thing where you get the AI to give you an output and then have some kind of specialist prompts or perhaps even specialist fine tune models come in and assess in various ways and then give it feedback and then let it reason some more and try to improve on what it just did. And we do see some of that stuff. I just talked to some OpenAI forward deployed engineers who are basically using that process to climb the hill on filing accuracy and that seems to be going quite well for them. So what techniques are you finding to be in practice most effective today? Yeah, so first we still do a bit of fine tuning. We still, I think it's at this point more of a kind of technique to make things have reasonable efficiency properties at scale than something to get the models to like new behaviors that you couldn't otherwise elicit. That said, so we do a lot of the, otherwise, we do a lot of the things that you pointed out. Let me maybe talk about one of them that we've been investing more in lately, which is what you could call like world models or like knowledge representations that make the models work more checkable. So they are, what is the motivation? So the motivation is actually, there's like maybe similar to the kind of medical case you had, which is like I had a friend who also had a cancer and it was like a case where I then used elicit to get a lot of the raw data for it. So I ran the systematic literature review flow and like for a few versions of the kind of question of like how do you address this particular type of cancer? And I did end up just with, even after filtering down all the information with a ton of papers, it was maybe after filtering out just for the most relevant papers, maybe it was still like 5,000 papers or so. And then the question is like, what do you do with that, right? Like you're now, that's, you could try to like somehow throw it all into like a million context, context window, but I don't think it would actually work that well and I think the model wouldn't be that good at coherently reasoning about it. And so then the question is like what else can you do? I think one thing that people have tried, I don't know if you're familiar with the Karpathy termed it like LLM Biki I think or something like that, you know, which is like you build this knowledge base as like sort of like a obsidian like knowledge base or like a folder of markdown files where you try, you tell your model, hey, we're researching this topic, organize the information in a way that makes sense and do these like iterative updates to it, like maybe it's a GitHub repo and the model gets to add new nodes to it, gets to move information from one file to another, gets to propagate information. And so I think that is a really interesting direction because if you think about how do we get to models that coherently answer complex questions, well, either it happens in the weights of the model or it happens in some explicit representation. And then what is the explicit representation? One, I think, text files are appealing. It's a nice start. But then you ask yourself, it's also very flexible. What properties do you want this representation to have such that it actually helps with the research questions you have? And once you think about that and you're like, well, I want it to let me make predictions about what's going to happen in my case. I want it to help me talk about like interventions.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence