High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Nathan Labenz: uncertainty

13 Jun 2026 The Cognitive Revolution AI in the AM — Week 2 Highlights (June 2026)

“Is it esoteric? I don't know. It strikes me as fairly important from the Fable system card that I'd love to get your take on.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
uncertainty
Recorded
13 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…So I think there's going to be a couple of different pieces of the story and different labs kind of emphasize different pieces to different extents. And so I think one piece, as you said earlier, is just monitoring, like look at them very carefully as they're doing things. It's very fundamental that monitoring of this form, if you do like chain of thought monitoring or white box monitoring or the like, that only takes you so far. And so then you need some story once that falls down, because you kind of go up the ramp. One of those next stories is, well, the models will find some other technique. They'll find kind of another solution to alignment which scales further. So that's sort of automated alignment of various kinds. But then I think there's other stories. So in some sense, all of the labs in various ways are doing some form of scalable oversight. And so they're getting models to supervise themselves. If you get that kind of, if you tie that knot correctly, then that could potentially scale very far, although there's various kind of known obstacles to that, which are not very well addressed. And then I think finally, there's this whole area of sort of character training and personas and so on, where they're trying to, in kind of, in kind of colloquially intervene on the models to be a good, to have good values such that, especially as you do this sort of like scale up oversight extrapolation, the good values preserve across that jump. And I think it's not, whether that will work instead of, there are fuzzy arguments why it could work. I think it's possible it will. We just don't understand that combination very well. And again, like a lot of the story is sort of monitoring, skill oversight, kind of character training, getting you far enough that you get into the automated alignment working regime and they find some better solution from the models. And I think that I would like to just push on all of those, because that's basically some kind of mad race, as you say, between the things we don't understand very well, but kind of are working pragmatically right now and the model's getting strong enough to blow through those. And I want to have some combination to make the prosaic things stronger or bring the automation automated solutions that give you stronger methods earlier. A mad race with monitoring carrying most of the weight, which is exactly where Prince took Friday's conversation when I raised the Fable system card. Is it esoteric? I don't know. It strikes me as fairly important from the Fable system card that I'd love to get your take on. And now we're getting these chains of thought that they show where it's just like lots of emojis. They call it illegible reasoning. They say this is an extreme example. But it's like it is indeed a pretty extreme example. I've been kind of struck in general by how much of the plan for recursive self-improvement seems to be monitoring in one way, shape, or form. You could dress that up and call it scalable oversight, but scalable oversight, as far as I can tell, is mostly a bunch of different angles on monitoring. How worried would you be or how much of an update do you think it is to see these extreme examples of illegible reasoning? Fantastic question. So I will say that, of course, I'm not an AI researcher, right? So this is going to be a deeply non-technical take for which I apologize in advance. So you're right. Like we've seen this. I think we've seen this for a while now and with OpenAI's models too. And so not a new phenon. view of the chain of thought is that it doesn't always reflect what the model is actually doing. But you do see these weird artifacts in the chain of thought and you kind of don't know what to do with them. I think what that teaches us is that monitoring just the chain of thought is probably not a perfect tool. Probably monitoring super intelligence generally is not a perfect tool. Because if a super intelligence knows what you're monitoring it, even if you can see its chain of thought and it's very legible to you, it can perhaps try to decide what to think so that you don't get alarmed. Right. And this is a lawyer's take, by the way, right? Like there's so many ways to phrase a particular thing that can be I guess there are ways to spin a particular thought, right, in different ways. Like if I have, if I'm gathering mushrooms and I've gathered 35 mushrooms and last week I gathered 20 mushrooms and what I need is 50, right? I can say, well, the number of mushrooms I've gathered has grown by almost 100%, which is great. Or I can say, well, I'm nowhere near 50. I'm so far behind, right? It's the same fact. So I don't know. I think that this problem of alignment and the risks are just there. And in my mind, there are certainly risks that the models will be thinking things that we don't know about. What does this all mean? It's hard to say. I think that we are tumbling into this future that will have probably super intelligence very fast. And In my view, there's no way to kind of stop it. So we need to be cognizant of these risks, try to monitor them as well as we can and take whichever actions are appropriate if we see something bad happening. But there's kind of no way, no conclusions to be drawn, right? No conclusions to be drawn other than yes, we should continue paying attention.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence