Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
26 Apr 2026 The Cognitive Revolution AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute
“It, it strikes me that we haven't really seen the true unleashing of the Internet's adversarial potential. And so, you know, that's one thing that they, I, I would say one of their biggest weaknesses, even Frontier models biggest weaknesses these days is how gullible they remain.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Apr 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, yeah. What's the NVIDIA model that you are kind of incorporating partnering with? Was it specifically trained to excel in the relevant search skills or it's just straight off the shelf? Is there anything that, I mean, do you envision this becoming something that will happen? I'm always personally a little wary of using small models because I just don't know what quality to expect and I don't want to find out the hard way. But I can easily imagine that one that is specifically trained to be a really good searcher would become competitive or even exceed what the frontier models would do, especially if it can take advantage of just extreme volume. Then I also do, especially because of Prakash's SEO question and your and your comments, I do wonder about adversarial robustness. It, it strikes me that we haven't really seen the true unleashing of the Internet's adversarial potential. And so, you know, that's one thing that they, I, I would say one of their biggest weaknesses, even Frontier models biggest weaknesses these days is how gullible they remain. So yeah, kind of curious what you think the training and specialization will look like as we go forward. Yeah, the Nematron model had just been released right before GTC, so it was not trained, especially in tool calling. It was a generalized model being trained on the various benchmarks. The models that are small and are more, you know, long lasting in the market are exceptionally good at at tool calling search. So Grok 41 fast is great at coming up with a set of queries. And of course the frontier models like, you know, Anthropic, you can see the how it calls, but you you can't really use a a frontier model for that like thinking model and and firing off other threads because it'll just slow down the overall experience. So yeah, so generally we use a smaller model and they're getting better all the time. I think people know that small models to to your worry whether small models are good. Everyone's talking about Cloud Code, right? The they use the Haiku models and so they reassured me the other day, oh, don't worry, I'm going to do this task with an LLM. But don't worry, I'm going to use Haiku. It's only $0.25 per million tokens, so that winds up to be less than $0.10 per thousand queries if we were trying to compare apples to apples. So I think our overall thought is that, you know, search can't be more expensive than intelligence. And then Haiku is being used for these high fidelity experiences of, of coding that you know, small models are more effective than than people are led to believe. VCX, by Fundrise, is the public ticker for private tech, giving everyday investors access to high-growth private companies in AI, space, defense tech, and more. Learn how to invest at https://getvcx.com Build your own Cognitive Revolution monitoring agent in one click.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.