Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
27 Jun 2026 The Cognitive Revolution AI:AM #4: Cameron on Model Consciousness, Duvenaud's Gradual Disempowerment, swyx's AI-Eng Alpha
“I think Even when the models have been, even when I said, six months ago, we were more just use the frontier to solve all the problems than we are today when we're exploring more open source, there still is lots of routing going on.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 27 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…My theory is for the application layer, we may have a sort of phase change moment, and it'll probably, exactly when this phase change kicks in, will vary by domain. Science will probably be one of the higher bars, where until you get to that bar, a strong default would be use the very best model, because you're dealing with scientists after all, they're going to want good output. And then at some point, as enough models cross a threshold, then your kind of strategic position is really about routing and kind of figuring out which is best for which thing, which is most cost effective for which thing. And that's one thing that presumably the frontier models will never do. I mean, you can kind of tell Claude, like, hey, sometimes you should delegate to Codex or whatever, and you know, I've set that up. But it's never going to be, kind of working fundamentally day and night in my interest to optimize in that way in the same way that you can do that for your customers. So how do you feel about that framing and where do you think we are with respect to, you know, just use Fable for everything versus, you know, you're actually creating strategic value by kind of playing this routing layer that helps people get better cost basis, but also, kind of avoid lock-in. just use Fable for everything versus, you know, you're actually creating strategic value by kind of playing this routing layer that helps people get better cost basis, but also, kind of avoid lock-in. Yeah, it's a great question. I think Even when the models have been, even when I said, six months ago, we were more just use the frontier to solve all the problems than we are today when we're exploring more open source, there still is lots of routing going on. And there's still lots of small tasks that we offload to very small, specialized, you know, sometimes even sub-billion parameter models. So even when that gap was so big, we still were doing routing. Something as simple as like, classify the field of study of the domain that the user's asking about to know how much we should care in certain search ranking, you know, what variables in search ranking we should care about the papers, right? Like, you know, in biomed, experimental design is incredibly, you should care so much about what was the sample size, what was the duration, where did the study take place, whereas computer science, that isn't a concept, and you should care much more about the recency and the citation velocity of the, or who the who the researchers were. To know how to care about each of those variables, we have a little model that will classify the field of study of the query. That does not need to be jammed into a giant prompt. That does not need to be a one-second latency API call to any frontier model. That should be a self-hosted, 800 million parameter model that you give a few fine-tuning examples. There's probably 20 different versions of those small little classification routers that inform downstream things that happen on query time. all of those should always be used by a really small model. Like it's just not even on the cost side. Like that's a very small amount of context you feed into it in a very small prompt needed just on the pure latency side. We can return those in sub.1 seconds in some cases. Base models are you using for that? Are you distilling from like a Claude into a what at that low scale to get those? And how much of the sort of frontier performance can you recover. if you take, for example, a liquid foundation model or whatever, you tell me what it is, or if you're willing to tell me what it is, you tell me what it is, and you're distilling into it, can you get, you know, in that narrow domain, 90% of the way back to Claude? Or like, how, what is the kind of Pareto curve look like when you are distilling the best into the fastest? How much performance can you retain? , in that narrow domain, 90% of the way back to Claude? Or like, how, what is the kind of Pareto curve look like when you are distilling the best into the fastest? How much performance can you retain? We've actually used human labels for some of those. We'll hire people to create very small data sets to fine-tune for those tasks. We'll also use models to sometimes create labels for those tasks where it's a... I guess you can call it a distillation process, but the objective isn't to get all of the representation of the weights of the entire model to do all these things. We're really trained to do a very small specialized thing. And then because of that answer, you can retain a ton of the performance, but it all depends on how specialized and how complex the task is. Like for a narrow, you know, 10 classification thing or all 10 classes of classification, it's all it's doing is that. But to give you something tangible to hold on to, I'd say you can get 95% of the performance from a frontier model for a very small classification task going all the way down to about a billion primer model. If you put in the work to give it a good fine tuning set. Yeah, very helpful. I appreciate the specificity.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.