High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

9 Jul 2026 The Cognitive Revolution AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen

“They basically think that what they're doing is somewhat near optimal and any sort of accuracy improvements you're going to get over them is going to be tiny and like hard to understand. And I think that's just because we only really understand human intelligence.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
9 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…ion. But if we were talking about some sort of post AGI world, what world would we be in with transformative AI? He thinks he would have a huge edge over the A is for exactly the reason that you gave. They are trained to try to predict things that have actually happened, and when things get wonky, you need some kind of creative lateral thinking. I think the rate of AI improvement is so astounding that I think that even the kind of lateral thinking, like trying to imagine a completely different scenario, will fall to the AIS. One unfortunate thing about it is it's hard to test this. So I think the more that AI continues doing strange things to the world and we wake up and see strange things in the news and those strange things are metaculous questions and on forecast bench and see teams like mine are trying to predict them better, we will actually get more evidence. But if there is a kind of a step change in the nature of the world, if we enter some sort of AGI transformative AI type of world, you know, we've got these geniuses and data centers, as people say, or anything like AI 2027 happens, then I think it's it's kind of going to be Wild West. I will say, I don't think humans are doing particularly great at imagining transformative AI. So the bar is a bit lower really. When you play with these AI forecasters, you will find them to be quite human and how they structure their reasoning. And again, this is not an accident. Like they're trained on how humans have structured their reasoning before. So a human forecaster would love to say, OK, what's the last 10 times something like this happened? What were the outcomes of those 10 times? And now I can make a distribution and say it's probably going to be something like this. The fact that an AI will do that, is it because it independently is arriving at the same conclusion? Is it because it's trained on humans doing that? Is it because it just thinks like a human? I don't think we have the answers to any of these questions now. Suffice to say, the superhuman reasoning is something that it's pretty hard to measure. Like, would you know it if you saw it because you were discussing before this? Using Fable a bunch, I have found that the way that Fable explains things is a little bit alien to the way that I find Opus or GPD 55 explaining things. It's very concise, I would say. Like it's very the sentences are shorter and full of jargon. It feels like it's compressing more information into a sentence than humans normally do. And to me, this is starting to get this, this showed off. It's kind of showing from behind the mask. Like the alien intelligence is a little bit more alien now than it was a month ago. ns normally do. And to me, this is starting to get this, this showed off. It's kind of showing from behind the mask. Like the alien intelligence is a little bit more alien now than it was a month ago. And I don't think it would be a wild prediction to say that we should expect more things like that to happen as post training is becoming more specialised from the labs, the models are getting larger. So how will this manifest in the Super forecasting perspective? Maybe actually super forecasting is the way to look at it. If you're looking at a better code, code base, you might say, well, you know, John Carmack would have written this like, OK, it looks really great, but like a great human would have done this too. But if you look at really brilliantly reasoned strategy about like, if the administration does this, then what will the outcomes be? You might start to see something that looks a little bit alien to the way that any human has analyzed it. And that might be an indication that the AI is actually starting to really surpass the humans. Keep in mind, this conversation happened the day before the Workspace paper landed. We asked Dan whether the chain of thought we see is even where the real action happens, a question the paper partly answered the next morning. Dan's answer stands on its own. I think there is a lot of detail in reality that is far beyond the human mind to understand. And as you approach more sophisticated intelligence, you will start seeing a lot of patterns. And then the point of trying to produce, you know, voxel perfect weather 3 weeks in the future is further away than people think. I think there's quite a lot of room. Human super forecasters don't tend to agree with me on this. They basically think that what they're doing is somewhat near optimal and any sort of accuracy improvements you're going to get over them is going to be tiny and like hard to understand. And I think that's just because we only really understand human intelligence. And when you kind of just zoom out from an information theory perspective from like a Kolkomorov complexity, like just modelling the world as byte strings, the AI overlords will eventually start to figure out stuff that is totally beyond humans to notice. But there's no way to prove this. My sense is that we will start to see it over the next year as the AIS will just get more and more accurate compared to humans in a way that humans don't even really understand. You look at the rationales of the forecast, it's like 5 paragraphs of dense reasoning and then a surprising conclusion and it will just not really make sense. But it'll just turn out to be really accurate and we will start to like understand it less and less as time goes on. If you simply ask a human super forecaster to explain their reasoning, they cannot actually make it fully legible. There is a layer of intuitive judgment that kind of feels like deep learning. They just kind of look at a bunch of evidence the way that a chess grandmaster just looks at a position and just sees the right move and they cannot explain it just popped into their head. p learning. They just kind of look at a bunch of evidence the way that a chess grandmaster just looks at a position and just sees the right move and they cannot explain it just popped into their head. You know, the grandmaster throws the knight and it just lands on the right square. Somehow. That happens with humans already. It happens with AI super forecasting systems today. So I think there's just no reason a priori to think that reasoning would always be legible. There's going to be some layer of intuitive judgement to the extent that the words intuitive judgement are referring to something going on in a large language model. It just has to be that way. Whether it is very that way or a little bit that way, I think is really your question, Nathan. Like is it if I just read the reasoning traces and I read the rationales and I see the research that it did, is it like more or less what a human would have done? I can kind of see where it's coming from. Or is it kind of inscrutable in the way where it just kind of discovers some new pattern in the world that no one has ever seen before, where we get what is the level for which it's doing something for which we cannot follow it down the deep dark forest into its reasoning? Almost by definition, we can't really know what that would look like. We asked about the economics of running it. It costs about a dollar or two to make a frontier forecast. That number can get a lot higher and it can go a little bit lower, but I think that's what you could anchor it again if you just looked at the cost per input and output tokens for an LLN, that gives you like a rough sense of the amount of research that would be done. One of the core questions that Future Search has tackled again, I described earlier how our main frontier was just doing present day research for quite a while until we got good enough that we could use it to improve forecasting. One of the questions we asked there was can you just pour more, more tokens into a question to get a more accurate answer? Again, it doesn't have to be a forecasting question. If I just ask you what is the current state of this clinical trial right now, Just give me the most accurate answer to that that you can. Can I just pour more tokens into that and get a more accurate answer? Again, this was kind of studied as deep research, writing these like 15 page reports with 700 citations. That was giving you a longer answer. Was it giving you a better answer? It wasn't super clear, which is why we studied this. Forecasting gives us an opportunity to do some world modelling.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence