Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen
“This is very useful for us because we can evaluate things immediately. So when Fable came out, the first time the clawed Fable came out, we were able to evaluate it within 24 hours and it was the best single agent forecaster on our leaderboard.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 9 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…You look at the Fortune 500 Nike. What does Nike have to fear from Anthropic? You get all the physical businesses out of the way and what you're left with is the pure IP businesses, right, software production, maybe pharma. I'm not I'm not so sure the paperwork businesses, banking, paperwork compliance businesses, accounting, tax compliance, regulatory, all of these things which are paperwork businesses, right? Those are all of the businesses where you have IP or relationships built up over years where if you have an Anthropic go in and they read through your entire workflow and processing, they can basically absorb all of that into the model. So that is where I think the risk is. The Frontier Labs are also horrible at sales, right? They're not. You look at IBMIBM has like 70% of the staff are basically sales engineers and the sales engineers are there to basically help you implement, maintain, you know, do all the grunt work, etcetera. The labs are not doing that. The labs have decided, especially Anthropic has decided to do this very lean structure of having almost no people at all and just putting out the models. And then just saying to like these enterprise teams here, you can go ahead and use it or not use it. And I'm the CTO is on, I'm signing 9 figure deals on the Uber over here, right? Like that guy isn't going to come and jump on your customer sales calls and like, oh, you know, sure, sure, we'll help you do this. And you know, our team will address this like next week. That's not happening, right? So the level of like customer service that is expected for enterprise SAS is not being provided by the Frontier Labs and they're not in a position to provide it. That's why they started this whole FTE program, but the FDE program, people thought it was a sales engineers program. It's actually a program to extract data and work flows and implement them inside the models themselves. And, and that is kind of, that is what Palantir, what Alex Carp is alluding to. He's like you, you're the FD ES are coming in and they're not there to help you to absorb your work flows. And once you absorb your work flows, you won't have a business because we take it and, and it's true. It's absolutely true. You know, that's what you know, we, we spoke to two opening IFDES and they went in into a company that Thrive owned rather than an external company. And as they went in then they took apart the workflow and they're basically absorbing it. And they said that the intention is to absorb it in the next round. And that is happening right now. Monday's guest Dan Schwartz, CEO of Future Search, 15 years in forecasting, formerly Metaculous, CTO and builder of Google's internal prediction market. Four days earlier, Scott Alexander had declared quote the AI super forecasters are here. Future searches systems now outscore the human super forecaster median on forecast bench. But how do you evaluate A forecaster without waiting months for the future to arrive? Their answer is called past casting. So the main thing about forecasting that's held it back. And again, this applies to human forecasting as well, as you generally have to wait for the future to happen to figure out if you were right. And with humans, they generally do this in year long tournaments. And so when the tournament ends, you find out which humans were best one year ago. Now humans don't get that much better over the course of one year. So finding out which humans were best one year ago is a very good indication of which who are the best humans today and how good objectively they are. This does not work with AI. If you wait one year and you find out who was good one year ago, you're getting a view of something very out of date. So one of the things Scott mentions in his article is that we used our best forecasting to predict stock returns. We published a set of stock rankings in August of 2025. It was basically a simple model for every stock based on forecasting certain fundamentals and extrapolating it out. We put it on the web, we pay you all the bit of it, and then we waited. And now it's been 10 months and that portfolio looks extremely good. But what is that really telling you? It's telling you that our forecasting in that particular methodology was good ten months ago, which is not something that most people will care about now. So we have a couple different forms of evidence. Some are more short term. There's tournaments running every couple of weeks, every couple of months. We at Future Search mostly rely on past casting. This is taking a snapshot of the Internet from some months ago and using the training window cut off of models to basically trick them into forecasting without the hindsight bias. This is very useful for us because we can evaluate things immediately. So when Fable came out, the first time the clawed Fable came out, we were able to evaluate it within 24 hours and it was the best single agent forecaster on our leaderboard. Everyone else had to wait weeks, weeks or months to find out how good clawed Fable actually was. So we internally, using the benchmark we call Bench to the Future, saw this progression kind of in real time. The rest of the world is seeing it kind of some months behind. And so if you if you read Stacott's article, you will see that over the last 12 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. st 12 months, the evidence has really come in. And over the last six months from these live forecasting tournaments and performance on actual prediction markets, you can see it's at least competitive. AI is competitive with humans and even teams of humans working together. Whether it's better is you got to synthesize a whole bunch of different disparate sources of evidence. I would say if you're curious about this or if you have forecasting needs in your life, you really should try it. So just go to Future Search. You get $20 free, so you can try some Frontier Forecast immediately. And I think you should judge for yourself whether you think they're good. We asked what the Frontier Labs should do with the forecaster this good. Yeah. So there's kind of two questions to this. One is what should they be doing with forecasting as a capability and what should they be doing with forecasting as an eval? So forecasting as a capability as kind of a business decision. What does say Open AI care whether ChatGPT is a good forecaster? I think that question is based on whether they, their consumers care about it. As a good forecaster. If you're entropic, I think you probably care more about the enterprise case. Like when people are using CLOD to do white collar work, do they care how good it is as a forecaster? Are people trying to use CLOD to make, say, financial forecast in an Excel spreadsheet? Is that something they care about?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.