High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Nathan Labenz: belief

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“I think human experts, doctors, lawyers, engineers, financiers, whoever, who are trying to make evals to try to produce the data for the frontier labs are finding that they are not smarter than the things being trained anymore.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
belief
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…st 12 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. AI is competitive with humans and even teams of humans working together. Whether it's better is you got to synthesize a whole bunch of different disparate sources of evidence. I would say if you're curious about this or if you have forecasting needs in your life, you really should try it. So just go to Future Search. You get $20 free, so you can try some Frontier Forecast immediately. And I think you should judge for yourself whether you think they're good. We asked what the Frontier Labs should do with the forecaster this good. Yeah. So there's kind of two questions to this. One is what should they be doing with forecasting as a capability and what should they be doing with forecasting as an eval? So forecasting as a capability as kind of a business decision. What does say Open AI care whether ChatGPT is a good forecaster? I think that question is based on whether they, their consumers care about it. As a good forecaster. If you're entropic, I think you probably care more about the enterprise case. Like when people are using CLOD to do white collar work, do they care how good it is as a forecaster? Are people trying to use CLOD to make, say, financial forecast in an Excel spreadsheet? Is that something they care about? ng CLOD to do white collar work, do they care how good it is as a forecaster? Are people trying to use CLOD to make, say, financial forecast in an Excel spreadsheet? Is that something they care about? So that's a business decision, and I can't really weigh in on that. I think, again, people will be discovering over time just how important forecasting is in everything, but it's going to be a slow process for humans to notice that. I think from an eval side, it's very different. Forecasting has this beautiful property that you basically get ground truth by waiting. So if I ask some question about the future and it basically an impossibly hard question, a question that even an AGI, an Oracle, a God could never really say because of chaos theory. Imagine just trying to predict, you know, like cubic meter weather 3 weeks in the future. Like you'd never be able to do it. But if you just wait, then you will see what that weather was in that cubic meter three weeks in the future. And so you basically have a completely limitless set of extremely hard, basically impossible questions where you get exact ground truth and there is no other e-mail like this. There is, if you want to say improve a coding harness, you just need to have more and more hard coding problems that are not in the training data for which you can say this is definitely the correct answer so that you can do some training on it. And that's hard. I think human experts, doctors, lawyers, engineers, financiers, whoever, who are trying to make evals to try to produce the data for the frontier labs are finding that they are not smarter than the things being trained anymore. And so if you can produce something that has a correct answer, the model's already going to figure out that correct answer. You need something where there's a correct answer and the model can't figure it out. Forecasting, I think, is the only completely and utterly renewable source of this. And again, this kind of is connected to forecasting is the the kind of Elon Musk's tweets, the quips, the forecasting is like the ultimate measure of intelligence. If you zoom out and think about it in one perspective, it is, again, I think coding intelligence, AI R&D intelligence, interpersonal intelligence are pretty darn important. I wouldn't say that forecasting is truly ultimate intelligence, but it is to some degree the ultimate eval. And I think this is something that Frontier labs like Canon should be paying attention to. One of the concerns a lot of people have about AI super forecasting is that it's 2IN distribution. I actually heard this from one of the very best forecasters I've ever had the pleasure of working with in my career. He basically said he believes that a system like Future Search would beat him head to head in a forecasting tournament about kind of near term outcomes of things that are within distribution. But if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the A is for exactly the reason that you gave. ion. But if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the A is for exactly the reason that you gave. They are trained to try to predict things that have actually happened, and when things get wonky, you need some kind of creative lateral thinking. I think the rate of AI improvement is so astounding that I think that even the kind of lateral thinking, like trying to imagine a completely different scenario, will fall to the AIS. One unfortunate thing about it is it's hard to test this. So I think the more that AI continues doing strange things to the world and we wake up and see strange things in the news and those strange things are metaculous questions and on forecast bench and see teams like mine are trying to predict them better, we will actually get more evidence. But if there is a kind of a step change in the nature of the world, if we enter some sort of AGI transformative AI type of world, you know, we've got these geniuses and data centers, as people say, or anything like AI 2027 happens, then I think it's it's kind of going to be Wild West. I will say, I don't think humans are doing particularly great at imagining transformative AI. So the bar is a bit lower really. When you play with these AI forecasters, you will find them to be quite human and how they structure their reasoning. And again, this is not an accident. Like they're trained on how humans have structured their reasoning before. So a human forecaster would love to say, OK, what's the last 10 times something like this happened? What were the outcomes of those 10 times? And now I can make a distribution and say it's probably going to be something like this. The fact that an AI will do that, is it because it independently is arriving at the same conclusion? Is it because it's trained on humans doing that? Is it because it just thinks like a human? I don't think we have the answers to any of these questions now. Suffice to say, the superhuman reasoning is something that it's pretty hard to measure. Like, would you know it if you saw it because you were discussing before this? Using Fable a bunch, I have found that the way that Fable explains things is a little bit alien to the way that I find Opus or GPD 55 explaining things. It's very concise, I would say. Like it's very the sentences are shorter and full of jargon. It feels like it's compressing more information into a sentence than humans normally do. And to me, this is starting to get this, this showed off. It's kind of showing from behind the mask. Like the alien intelligence is a little bit more alien now than it was a month ago.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence