High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nathan Labenz: evaluation

21 Jun 2026 The Cognitive Revolution AI:AM #3: Zvi on Fable, the Cases For & Against the Ban, + AI for Math, Logistics & More

“The other thing that's kind of related to this that jumped out at me is a sort of escalation, I guess, of both the difficulty of monitoring and some recent advances in monitoring techniques that I'm not sure exactly where they leave us on net. But we both see in the system card examples of extremely illegible chain of thought, which, you know, is just like this wall of emojis and sort of, you know, non-human language symbols strung together that I think is pretty spooky and, like, definitely, um, you know, don't like to see that, to put it simply and mildly.”

— Nathan Labenz

Source trail

Everything needed to verify it.

Speaker
Nathan Labenz
Attribution
Verified speaker
Claim type
evaluation
Recorded
21 Jun 2026
Publisher
The Cognitive Revolution

Transcript context

…ds that in fact are willing to, like, respond to how they are expected to be treated and are treated to end up being able to coordinate in reasonable ways. You can also... This also applies acausally. So, like, you have to consider the implications of your decision not only on other minds that exist now but other minds that existed in the past and will exist in the future. So to the extent that they are coordinate, they are correlated with us and that they're, these reactions are all intertwined, this can cause them to potentially treat us well, even if there is no direct current reason for them to treat us well. And that is also very helpful. But again, like, this is super complicated and, like, not today. Two more findings, both about whether we can even read what these models are thinking. An increasingly illegible chain of thought and a new interpretability tool that caught Fable doing something it never said out loud. Yeah, sort of preview of the global brain there in your comments. The other thing that's kind of related to this that jumped out at me is a sort of escalation, I guess, of both the difficulty of monitoring and some recent advances in monitoring techniques that I'm not sure exactly where they leave us on net. But we both see in the system card examples of extremely illegible chain of thought, which, you know, is just like this wall of emojis and sort of, you know, non-human language symbols strung together that I think is pretty spooky and, like, definitely, um, you know, don't like to see that, to put it simply and mildly. And then at the same time, we also have the natural language autoencoder work from Anthropic, which, and again, I'll assume folks are familiar with sparse autoencoders, basically a similar concept, except instead of creating a sparse feature-by-feature representation, you are actually creating a natural language representation of what the model is thinking at that given time. And from that natural language bottleneck, the autoencoder has to then feedback signal that allows the model to succeed in the way that it was originally going to succeed on the task. So hopefully this is faithful. Hopefully it's human readable. It seems like it's working pretty well so far. There are examples of the model knowingly, which I, I think you're right to really emphasize that distinction. It's like one thing if it's making a mistake and doing something bad because it thinks it was good or it didn't realize it was bad or whatever, but it's another thing really to, to zero in on if it knows it's doing bad and is ex- going ahead and doing it anyway. And so we do see things like in the natural language representation from the natural language autoencoder, things where it's, like, not supposed to access the internet, but it goes ahead to, goes ahead and tries to do it. And the natural language autoencoder representation is something like string concatenation trick to bypass URL filter. So it's clear that it understands that there's a filter and it's coming up with a trick to work around it. And so that's not good. But then, you know, we do see that the technique is able to surface that from model internals without necessarily having it verbalized in the chain of thought, which is good because, again, those are getting, like, at least in some instances, um, quite opaque. And then the safety classifiers wrapped around all of it. Why, Zvi argues, they can only survive by being almost comically blunt, and what that tells you about the difference between defending against a person and defending against a mind. But the classifiers... So it's much easier to think about a pink elephant than to not think about a pink elephant, right? Even though most of the time you succeed at not thinking about a pink elephant, almost always actually. To consciously decide not to do so is often hard, but consciously doing so is really easy. So it is very possible that classifiers can survive as long as they're willing to endure false positives. Like, the classifiers in Fable have ludicrous amounts of false positives, right? Like, you say the word cancer and you get cut off. Like, just levels of false positive. But that's intentional because, like, they're not even necessarily false positives because people think of it as the false positive is I wasn't trying to create a bioweapon. We know that. You were trying to talk about biology, and we've decided that no, this model just doesn't talk about biology at all. Like, you know, it's not that we don't talk about what Bruno sees, it's we don't talk about Bruno, period. Bruno does not exist, right? And so they're like, "This is a false positive. He's just my brother." Like, we don't talk about Bruno. We don't talk about Bruno. And so- The classifiers seem like they actually succeeded. It's just that they chose a giant blast radius because of the adversarial problems, basically.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence