High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“Forecasting gives us an opportunity to do some world modelling. So Future Search talked about this a little bit at the Manifest conference a couple weeks ago and the feature in the product is rolling out I think literally today.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…p learning. They just kind of look at a bunch of evidence the way that a chess grandmaster just looks at a position and just sees the right move and they cannot explain it just popped into their head. You know, the grandmaster throws the knight and it just lands on the right square. Somehow. That happens with humans already. It happens with AI super forecasting systems today. So I think there's just no reason a priori to think that reasoning would always be legible. There's going to be some layer of intuitive judgement to the extent that the words intuitive judgement are referring to something going on in a large language model. It just has to be that way. Whether it is very that way or a little bit that way, I think is really your question, Nathan. Like is it if I just read the reasoning traces and I read the rationales and I see the research that it did, is it like more or less what a human would have done? I can kind of see where it's coming from. Or is it kind of inscrutable in the way where it just kind of discovers some new pattern in the world that no one has ever seen before, where we get what is the level for which it's doing something for which we cannot follow it down the deep dark forest into its reasoning? Almost by definition, we can't really know what that would look like. We asked about the economics of running it. It costs about a dollar or two to make a frontier forecast. That number can get a lot higher and it can go a little bit lower, but I think that's what you could anchor it again if you just looked at the cost per input and output tokens for an LLN, that gives you like a rough sense of the amount of research that would be done. One of the core questions that Future Search has tackled again, I described earlier how our main frontier was just doing present day research for quite a while until we got good enough that we could use it to improve forecasting. One of the questions we asked there was can you just pour more, more tokens into a question to get a more accurate answer? Again, it doesn't have to be a forecasting question. If I just ask you what is the current state of this clinical trial right now, Just give me the most accurate answer to that that you can. Can I just pour more tokens into that and get a more accurate answer? Again, this was kind of studied as deep research, writing these like 15 page reports with 700 citations. That was giving you a longer answer. Was it giving you a better answer? It wasn't super clear, which is why we studied this. Forecasting gives us an opportunity to do some world modelling. citations. That was giving you a longer answer. Was it giving you a better answer? It wasn't super clear, which is why we studied this. Forecasting gives us an opportunity to do some world modelling. So Future Search talked about this a little bit at the Manifest conference a couple weeks ago and the feature in the product is rolling out I think literally today. The idea is that once you have a repository of forecasts, every marginal forecast can draw on the implicit world model in those forecasts in order to give you a better answer. Co founder of Future Search, Lawrence Phillips wrote this up on Less Wrong a couple months ago and it was a bit neglected. He basically made the case that as a public good, if you produce this kind of large body of forecasting questions that kind of fed into each other and remain mutually consistent, you could understand the world dramatically better. And the main barrier to that is simply when you put more tokens into your world model effectively, does it get better or does it get worse? And I think his big insight was around January or February around Opus 4.6 sometime or GPD 5.4 somewhere around there. For the very first time, it became possible to put more tokens into like a broad research thing and actually get a better answer. Not one that just trails off into nonsense kind of garbage in, garbage out. And Future Search is doing this in its products. And that's the other reason that we have a consumer product is because the more people that forecast, the better the forecast will be for them. And then in theory, the better the forecast will be for everybody as we build this deeper implicit model of the world. Now, many companies and research labs have had these ideas of building world models. Again, world model, the way I I use that term is maybe misleading. A lot of people talk about like geospatial reasoning, like I'm trying to build a robot hand that can go and pick something up. That's a world model as well. I mean, it's a world model of like what is going on in the world that helps me predict outcomes in a very basic way. So again, more broadly, I think the big question is, can you just point more tokens into more research and get better research of any kind? Again, AI research, coding, like whatever you know. Dan mentioned Future Search is building what they call a world model. Thousands of mutually consistent forecasts, each new one drawing on all the others. I asked what structure that actually takes because I'd seen this movie before. What is the structure that ultimately gets instantiated into? Like, are we talking about a graph database? I feel like those kinds of ideas make sense for this sort of thing, but I also could imagine that they might introduce some weird failure modes and I I guess in general I'm I'm so there's the how does it get instantiated question, then this other question that's kind of in the back of my mind. Fun fact about me, I participated. I was actually on the good judgement team way back in the like DARPA forecasting challenge. Or was it IR by for whoever funded that 15 plus years ago. I did well, but not like top, top tier super forecaster. ctually on the good judgement team way back in the like DARPA forecasting challenge. Or was it IR by for whoever funded that 15 plus years ago. I did well, but not like top, top tier super forecaster. And at the same time I also worked briefly at a financial services consulting firm that had done a lot of the financial risk modeling for Fannie Mae. And I don't probably have to tell you how that story turned out, but there was a lot of expert forecasting that was instantiated in this very spreadsheet kind of causal graph sort of way, right? Where all these you could literally like, you know, hit the sort of the one visualization button in Excel and you'd see these like colored arrows, you know, fanning out from cell to cell. And somehow in the end, it was just all totally off. So I do wonder about how you think about like correlated failures as you build out these world models or if there's any kind of correction mechanism or something to say, you know, wait a second, What happens if we do have some house, housing prices never go down nationwide, kind of bad assumption lurking in our world model. Is there a way to detect that? Obviously humans have this problem too, right? The financial crisis proves that. But you can imagine the next one being even way worse, right? Because we're like very reliant on a very small set of AI minds that are, you know, working at it from 1000 different directions. But they may have somewhat consistent flaws in their reasoning as they go. Can we protect ourselves against that in in any way? Definitely we can. I will try to answer that both theoretically and with an anecdote. So I tried to world model the Fable situation when it got banned because I wanted it, but also it was kind of a good forecasting question and there was some nice money trading on Kolchi and Polymarket. And I made exactly the mistake that you're talking about, Nathan. So I ran a bunch of future search forecasts and I kind of just manually went through them. There was a couple of scenarios. Some conditional forecasts are basically 33 load bearing forecasts, basically starting from what even happened. Like why did the government issue this export control? Was it a simple misunderstanding? Is this political leverage?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence