Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“I would say if you're curious about this or if you have forecasting needs in your life, you really should try it.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Monday's guest Dan Schwartz, CEO of Future Search, 15 years in forecasting, formerly Metaculous, CTO and builder of Google's internal prediction market. Four days earlier, Scott Alexander had declared quote the AI super forecasters are here. Future searches systems now outscore the human super forecaster median on forecast bench. But how do you evaluate A forecaster without waiting months for the future to arrive? Their answer is called past casting. So the main thing about forecasting that's held it back. And again, this applies to human forecasting as well, as you generally have to wait for the future to happen to figure out if you were right. And with humans, they generally do this in year long tournaments. And so when the tournament ends, you find out which humans were best one year ago. Now humans don't get that much better over the course of one year. So finding out which humans were best one year ago is a very good indication of which who are the best humans today and how good objectively they are. This does not work with AI. If you wait one year and you find out who was good one year ago, you're getting a view of something very out of date. So one of the things Scott mentions in his article is that we used our best forecasting to predict stock returns. We published a set of stock rankings in August of 2025. It was basically a simple model for every stock based on forecasting certain fundamentals and extrapolating it out. We put it on the web, we pay you all the bit of it, and then we waited. And now it's been 10 months and that portfolio looks extremely good. But what is that really telling you? It's telling you that our forecasting in that particular methodology was good ten months ago, which is not something that most people will care about now. So we have a couple different forms of evidence. Some are more short term. There's tournaments running every couple of weeks, every couple of months. We at Future Search mostly rely on past casting. This is taking a snapshot of the Internet from some months ago and using the training window cut off of models to basically trick them into forecasting without the hindsight bias. This is very useful for us because we can evaluate things immediately. So when Fable came out, the first time the clawed Fable came out, we were able to evaluate it within 24 hours and it was the best single agent forecaster on our leaderboard. Everyone else had to wait weeks, weeks or months to find out how good clawed Fable actually was. So we internally, using the benchmark we call Bench to the Future, saw this progression kind of in real time. The rest of the world is seeing it kind of some months behind. And so if you if you read Stacott's article, you will see that over the last 12 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. st 12 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. AI is competitive with humans and even teams of humans working together. Whether it's better is you got to synthesize a whole bunch of different disparate sources of evidence. I would say if you're curious about this or if you have forecasting needs in your life, you really should try it. So just go to Future Search. You get $20 free, so you can try some Frontier Forecast immediately. And I think you should judge for yourself whether you think they're good. We asked what the Frontier Labs should do with the forecaster this good. Yeah. So there's kind of two questions to this. One is what should they be doing with forecasting as a capability and what should they be doing with forecasting as an eval? So forecasting as a capability as kind of a business decision. What does say Open AI care whether ChatGPT is a good forecaster? I think that question is based on whether they, their consumers care about it. As a good forecaster. If you're entropic, I think you probably care more about the enterprise case. Like when people are using CLOD to do white collar work, do they care how good it is as a forecaster? Are people trying to use CLOD to make, say, financial forecast in an Excel spreadsheet? Is that something they care about? ng CLOD to do white collar work, do they care how good it is as a forecaster? Are people trying to use CLOD to make, say, financial forecast in an Excel spreadsheet? Is that something they care about? So that's a business decision, and I can't really weigh in on that. I think, again, people will be discovering over time just how important forecasting is in everything, but it's going to be a slow process for humans to notice that. I think from an eval side, it's very different. Forecasting has this beautiful property that you basically get ground truth by waiting. So if I ask some question about the future and it basically an impossibly hard question, a question that even an AGI, an Oracle, a God could never really say because of chaos theory. Imagine just trying to predict, you know, like cubic meter weather 3 weeks in the future. Like you'd never be able to do it. But if you just wait, then you will see what that weather was in that cubic meter three weeks in the future. And so you basically have a completely limitless set of extremely hard, basically impossible questions where you get exact ground truth and there is no other e-mail like this. There is, if you want to say improve a coding harness, you just need to have more and more hard coding problems that are not in the training data for which you can say this is definitely the correct answer so that you can do some training on it. And that's hard. I think human experts, doctors, lawyers, engineers, financiers, whoever, who are trying to make evals to try to produce the data for the frontier labs are finding that they are not smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting is the the kind of Elon Musk's tweets, the quips, the forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is, again, I think coding intelligence, AI R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier labs like Canon should be paying attention to. One of the concerns a lot of people have about AI super forecasting is that it's 2IN distribution. I actually heard this from one of the very best forecasters I've ever had the pleasure of working with in my career. He basically said he believes that a system like Future Search would beat him head to head in a forecasting tournament about kind of near term outcomes of things that are within distribution. But if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the A is for exactly the reason that you gave.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.