High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Nathan Labenz: prediction

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“One unfortunate thing about it is it's hard to test this. So I think the more that AI continues doing strange things to the world and we wake up and see strange things in the news and those strange things are metaculous questions and on forecast bench and see teams like mine are trying to predict them better, we will actually get more evidence.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
prediction
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…ng CLOD to do white collar work, do they care how good it is as a forecaster? Are people trying to use CLOD to make, say, financial forecast in an Excel spreadsheet? Is that something they care about? So that's a business decision, and I can't really weigh in on that. I think, again, people will be discovering over time just how important forecasting is in everything, but it's going to be a slow process for humans to notice that. I think from an eval side, it's very different. Forecasting has this beautiful property that you basically get ground truth by waiting. So if I ask some question about the future and it basically an impossibly hard question, a question that even an AGI, an Oracle, a God could never really say because of chaos theory. Imagine just trying to predict, you know, like cubic meter weather 3 weeks in the future. Like you'd never be able to do it. But if you just wait, then you will see what that weather was in that cubic meter three weeks in the future. And so you basically have a completely limitless set of extremely hard, basically impossible questions where you get exact ground truth and there is no other e-mail like this. There is, if you want to say improve a coding harness, you just need to have more and more hard coding problems that are not in the training data for which you can say this is definitely the correct answer so that you can do some training on it. And that's hard. I think human experts, doctors, lawyers, engineers, financiers, whoever, who are trying to make evals to try to produce the data for the frontier labs are finding that they are not smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting is the the kind of Elon Musk's tweets, the quips, the forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is, again, I think coding intelligence, AI R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier labs like Canon should be paying attention to. One of the concerns a lot of people have about AI super forecasting is that it's 2IN distribution. I actually heard this from one of the very best forecasters I've ever had the pleasure of working with in my career. He basically said he believes that a system like Future Search would beat him head to head in a forecasting tournament about kind of near term outcomes of things that are within distribution. But if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the A is for exactly the reason that you gave. ion. But if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the A is for exactly the reason that you gave. They are trained to try to predict things that have actually happened, and when things get wonky, you need some kind of creative lateral thinking. I think the rate of AI improvement is so astounding that I think that even the kind of lateral thinking, like trying to imagine a completely different scenario, will fall to the AIS. One unfortunate thing about it is it's hard to test this. So I think the more that AI continues doing strange things to the world and we wake up and see strange things in the news and those strange things are metaculous questions and on forecast bench and see teams like mine are trying to predict them better, we will actually get more evidence. But if there is a kind of a step change in the nature of the world, if we enter some sort of AGI transformative AI type of world, you know, we've got these geniuses and data centers, as people say, or anything like AI 2027 happens, then I think it's it's kind of going to be Wild West. I will say, I don't think humans are doing particularly great at imagining transformative AI. So the bar is a bit lower really. When you play with these AI forecasters, you will find them to be quite human and how they structure their reasoning. And again, this is not an accident. Like they're trained on how humans have structured their reasoning before. So a human forecaster would love to say, OK, what's the last 10 times something like this happened? What were the outcomes of those 10 times? And now I can make a distribution and say it's probably going to be something like this. The fact that an AI will do that, is it because it independently is arriving at the same conclusion? Is it because it's trained on humans doing that? Is it because it just thinks like a human? I don't think we have the answers to any of these questions now. Suffice to say, the superhuman reasoning is something that it's pretty hard to measure. Like, would you know it if you saw it because you were discussing before this? Using Fable a bunch, I have found that the way that Fable explains things is a little bit alien to the way that I find Opus or GPD 55 explaining things. It's very concise, I would say. Like it's very the sentences are shorter and full of jargon. It feels like it's compressing more information into a sentence than humans normally do. And to me, this is starting to get this, this showed off. It's kind of showing from behind the mask. Like the alien intelligence is a little bit more alien now than it was a month ago. ns normally do. And to me, this is starting to get this, this showed off. It's kind of showing from behind the mask. Like the alien intelligence is a little bit more alien now than it was a month ago. And I don't think it would be a wild prediction to say that we should expect more things like that to happen as post training is becoming more specialised from the labs, the models are getting larger. So how will this manifest in the Super forecasting perspective? Maybe actually super forecasting is the way to look at it. If you're looking at a better code, code base, you might say, well, you know, John Carmack would have written this like, OK, it looks really great, but like a great human would have done this too. But if you look at really brilliantly reasoned strategy about like, if the administration does this, then what will the outcomes be? You might start to see something that looks a little bit alien to the way that any human has analyzed it. And that might be an indication that the AI is actually starting to really surpass the humans. Keep in mind, this conversation happened the day before the Workspace paper landed. We asked Dan whether the chain of thought we see is even where the real action happens, a question the paper partly answered the next morning. Dan's answer stands on its own. I think there is a lot of detail in reality that is far beyond the human mind to understand. And as you approach more sophisticated intelligence, you will start seeing a lot of patterns. And then the point of trying to produce, you know, voxel perfect weather 3 weeks in the future is further away than people think. I think there's quite a lot of room. Human super forecasters don't tend to agree with me on this. They basically think that what they're doing is somewhat near optimal and any sort of accuracy improvements you're going to get over them is going to be tiny and like hard to understand. And I think that's just because we only really understand human intelligence. And when you kind of just zoom out from an information theory perspective from like a Kolkomorov complexity, like just modelling the world as byte strings, the AI overlords will eventually start to figure out stuff that is totally beyond humans to notice. But there's no way to prove this. My sense is that we will start to see it over the next year as the AIS will just get more and more accurate compared to humans in a way that humans don't even really understand. You look at the rationales of the forecast, it's like 5 paragraphs of dense reasoning and then a surprising conclusion and it will just not really make sense. But it'll just turn out to be really accurate and we will start to like understand it less and less as time goes on. If you simply ask a human super forecaster to explain their reasoning, they cannot actually make it fully legible. There is a layer of intuitive judgment that kind of feels like deep learning. They just kind of look at a bunch of evidence the way that a chess grandmaster just looks at a position and just sees the right move and they cannot explain it just popped into their head.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence