High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“And I tried my best over basically like, you know, 12 to 16 hours of the Fable situation. I think I made a pretty good model.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…ctually on the good judgement team way back in the like DARPA forecasting challenge. Or was it IR by for whoever funded that 15 plus years ago. I did well, but not like top, top tier super forecaster. And at the same time I also worked briefly at a financial services consulting firm that had done a lot of the financial risk modeling for Fannie Mae. And I don't probably have to tell you how that story turned out, but there was a lot of expert forecasting that was instantiated in this very spreadsheet kind of causal graph sort of way, right? Where all these you could literally like, you know, hit the sort of the one visualization button in Excel and you'd see these like colored arrows, you know, fanning out from cell to cell. And somehow in the end, it was just all totally off. So I do wonder about how you think about like correlated failures as you build out these world models or if there's any kind of correction mechanism or something to say, you know, wait a second, What happens if we do have some house, housing prices never go down nationwide, kind of bad assumption lurking in our world model. Is there a way to detect that? Obviously humans have this problem too, right? The financial crisis proves that. But you can imagine the next one being even way worse, right? Because we're like very reliant on a very small set of AI minds that are, you know, working at it from 1000 different directions. But they may have somewhat consistent flaws in their reasoning as they go. Can we protect ourselves against that in in any way? Definitely we can. I will try to answer that both theoretically and with an anecdote. So I tried to world model the Fable situation when it got banned because I wanted it, but also it was kind of a good forecasting question and there was some nice money trading on Kolchi and Polymarket. And I made exactly the mistake that you're talking about, Nathan. So I ran a bunch of future search forecasts and I kind of just manually went through them. There was a couple of scenarios. Some conditional forecasts are basically 33 load bearing forecasts, basically starting from what even happened. Like why did the government issue this export control? Was it a simple misunderstanding? Is this political leverage? re basically 33 load bearing forecasts, basically starting from what even happened. Like why did the government issue this export control? Was it a simple misunderstanding? Is this political leverage? This is really about foreigner threat because Fable is actually dangerous for hacking, etcetera. We didn't know those things. So I kind of put it all together and when I looked at all of the outcomes and I talked about it with Claude Code a lot, one thing came out, which is basically every forecast and every scenario. I had thought that access would come to Americans 1st and then foreigners at some later point in the future, and that was wrong. When it came out last week, it came back for everybody, so clearly there was some wait in one of my scenarios that was wrong, but I had a basically like kind of a correlated failure in there somewhere. I still haven't completely understood where my reasoning was wrong. It's also possible I just got really unlucky and the outcome we're in was just extremely unlikely. There's an n = 1. You can never know if anyone forecast is great. That's one of the hard things about it. But I think I systematically got it wrong by having a bunch of correlated reasoning failures across my various scenarios, so this definitely does happen. Metaculous has a system like this. In the years since I was the CTO there, they have built an actual causal graph platform and product. So you can go to the Metaculous site and click right and you'll find it there. I think the field still generally believes that things like this will work, but nobody has actually made a good one before. And I tried my best over basically like, you know, 12 to 16 hours of the Fable situation. I think I made a pretty good model. I think I did. I was close to having a very accurate forecast, but I didn't quite get it. I don't think. I don't think those meticulous models on their website right now are so amazing, but I do fundamentally believe in the approach. As you're saying, Nathan, this has been tried for a long time. When I was the CTO of Metangolis, honestly, it was it was kind of the dream. It was the Holy Grail. Can we tie all of these forecasts together into some sort of causal graph? And I think what I can say is that AI makes this tractable. There was just no way that that was going to work with a bunch of human economists looking at Freddie Mac or Fannie Mae. I can totally understand why that method didn't work for them then. Whether AI can make it work right now is unclear. Whether AI will make this work in general feels nearly guaranteed, and I don't think Future Search is the only org that is working on this right now. Before he left the unhedged version. So Future Search contributed some forecast to AI 2027 and we studied that problem pretty seriously with the evidence of a little bit over a year ago. And we built a model of R&D take off speeds under the kind of the core AI 2027 scenario where the main way things get crazy is that AI is used more in the development of AI, first by achieving the superhuman coder milestone and then the superhuman AI researcher milestone. And I am unhappy to report that I think that story is generally correct. re in the development of AI, first by achieving the superhuman coder milestone and then the superhuman AI researcher milestone. And I am unhappy to report that I think that story is generally correct. I don't know if the timelines are exactly right, but I my forecast from that process of leading to something that looks like super intelligence around 2031 is roughly stable. I think the things that have happened in the years since AI 2027 come out very much indicate the theory that the most important thing going on is how useful is AI and improving the productivity of AI researchers within Frontier Labs. I've made public predictions that I thought Anthropic was going to run away with it because they had the best feedback loop of talent and actually using their AI internally. I think that has been, you know, n = 1. But I think it's been totally shown that that's been happening recently. So I think that will continue to happen. And Dan's closing confession about the whole project to prediction markets and a hope for what AI forecasting could still become. Maybe just in closing, sketch out a little bit more of the future as you hope it might unfold. Not necessarily the most likely scenario, because maybe the most likely thing is people act foolishly and don't take advantage of the benefits of forecasting. But like, if we really do a good job right, and we're and we're interested in truth seeking and we get the AIS working as well as you think they might, how do you think life feels different? Yeah, I have to leave with another example of me being a bad forecaster. I, I guess everyone who tries forecasting thinks they're a bad forecaster because they see things getting wrong. Here's a prediction that I made really strongly 5 or 10 years ago that has basically been totally falsified. I predicted that if we had highly visible, highly liquid prediction markets that were covering all of the like the major technological and political and economic things going on, that humanity would be wiser and people would make better decisions in government. So here we are. We have Pauline market and call sheet.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence