Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
21 Jun 2026 The Cognitive Revolution AI:AM #3: Zvi on Fable, the Cases For & Against the Ban, + AI for Math, Logistics & More
“I think that we should actually sort of sit here and, and frame what is actually happening when we say, like, Fable outperforms on frontier code.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 21 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…The problem is that most of the safety training is done in post-training, not in pre-training. So once the jailbroken model is there, uh, once the model is jailbroken, you can do whatever you want a lot of the time. And so we set out to try to solve that at an earlier stage. And one of the things that we've been accelerating is an approach called gradient routing, which we-- which basically winds up in pre-training, you route different dangerous capabilities into different experts in a mixture of experts models. So you wind up having some dangerous experts that learn specifically the CBRN stuff or the cyber stuff, and then you can later ablate those experts, and this winds up, uh, so you completely remove it. So you have the regular model, and then you have the safe model that, that, that winds up being public. And this is, this has been going decently well. It's like a, it's a, it's a, it's still an early stage and neglected approach, but we're excited to release it fairly soon because it potentially, it potentially solves this big issue that, uh, a lot of people are very concerned about right now today. And our larger thesis is that if the field had been investing more in AI alignment R&D instead of just scaling compute, if we'd done this earlier on, we would have found techniques like this, and you wouldn't have the issue right now, uh, with the Trump admin and, and Anthropic because this would be al-already in Fable 5. Then the software itself. Eno Reyes runs Factory, which builds the systems that build code, and his read on why Fable wins the big coding benchmark is the most honest thing I heard a builder say all week. It is not the answer you would expect. I think that we should actually sort of sit here and, and frame what is actually happening when we say, like, Fable outperforms on frontier code. So frontier code, good, great benchmark. Like, that's a-- I, I'm really glad that people like the cognition team are, like, thinking through how do we, how do we measure on more novel and difficult problems, like the types of challenges that contemporary models are facing. And so I think we need more of those. There's an- another great benchmark called Program Bench that also looks at, like, reverse engineering on extremely hard problems. The pass rate there is, like, effectively zero. Um, we have internal benchmarks that we have zero percent pass rates on. Uh, and, and I think that, like, generally this is great when the, when we introduce these new benchmarks. But if you think about what it means to score on a benchmark I mean, you can sort of read through, right? Oh, well, we assessed correctness by running tests. We used LLMs to judge correctness. We built novel verifiers specific to the problem. Basically, what that means is that when somebody spends forty-plus hours creating a verification of a single code change, we can then reliably evaluate if the model was good at working on that problem. That is, like, totally reasonable, but I think what it translates to is that in the real world, people-- the challenge is often not can the model write code that works? It's basically every other aspect like can I trust that this model output code that works? Um, does this model have the, the deterministic feedback loops inside of the code base to get to that correctness? The set of repositories in that benchmark are all very well-tested, very well-known open source code bases, right, where the maintainers approved it. The level of rigor of what we would call agent readiness in open source code bases actually tends to be much higher than in enterprises. And so-- Which makes sense. You're basically accepting changes from the outside world from random people. How different is that from coding agents where you're sort of, like, getting changes that you sort of lightly asked for and you don't even know the source? It's kind of black box generation, right? And so I think a lot of open source maintainers have gone through the, the rigor and the effort to add these deterministic verification and validation loops into their system such that when a new change comes in, you think about how did Fable get such a high score? Well, it ran the tests, it ran the linters, it did more focused application of the type checking. It used all of these tools to hill climb its way to high success. And I think that in general, the-- if you don't have those things, you're screwed no matter what. And what we would sort of argue is that all of these pieces are part of the puzzle. You can't just plop good model. You can't just have agent readiness with a bad model. Like, y-y-you sort of need to go through and invest in upgrading the basis by which your company has these feedback loops. You have to upgrade the way you think about this because it's a risk thing. Like, humans have to say, "I'm going to, at this point now, start accepting code changes that I haven't read." And then third, you do need great models. So, so I think that, like, basically Opus four point six maybe was-- has been sufficient enough. I would even argue that before then we've had models that were sufficient enough to go full auto. I think that all of these other things need to catch up in order to then take advantage of these models. And basically, the gains we see in models today are primarily coming from effectively, like, the models getting better at, at getting away with not using these verification loops like humans are. So there's this giant feedback loop that's extremely human-driven right now.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.