Evidence receipt / recommendation
Published · transcript-backedFlo Crivello: recommendation
14 Aug 2026 The Cognitive Revolution Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
“We have found you should almost never use multiple models powering the same agent, and that we have found one exception to this rule is when you spin up new blank sub agents.”
Source trail
Everything needed to verify it.
- Speaker
- Flo Crivello
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 14 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…where to integrate them, especially because it didn't sound easy before, but in a context where you have your workflows and there's nodes in the sort of graph of work, you could go into a particular node and say, okay. We kinda know what the inputs here are and the outputs, it's, like, relatively controlled environment and so we can do structured testing. I did all that and you still reported having some false positives over time where a model could pass a bunch of tests and you would feel good about it and then if you'd start to test it live with users, you'd get the feedback that like, hey. Like, it got dumb and I'm not you know? Somehow it was, like, still hard to measure even with, like, a much more structured environment for the AI to work in. Now as you're in this very open ended, just tag Lindy in Slack and send it anything, it seems like that problem would have increased in difficulty dramatically. So how are you approaching it, and what are we learning in terms of where the open source models and it's not even so much I think Anthropic does make some good points about this sometimes. Obviously, there's more to the story, but I do think they they're apt in some ways where they say, in some cases, it's less about open and closed source and more about just like what's capability level, what's the price, what are the the features. So regardless, I guess, of whether you're even going to open source or just going to Haiku, how do you decide when you can do it? Great question. We have found you should almost never use multiple models powering the same agent, and that we have found one exception to this rule is when you spin up new blank sub agents. Emphasis on blank because there's two types of sub agents. You have a blank sub agent, then you have a full sub agent, which inherits the context window of the of the parent agent. And the reason why you don't wanna do that is because of caching. We we we just obsess about caching, obviously, because it's just so expensive otherwise. Like, the economics do not they barely work with caching. They just cannot work without caching. So caching is is a must have. And and so I'll give you an example, the validator system that I mentioned. So it's this LLMS judge. By the way, highly recommend anyone who builds AI agents. Like, this is one of the lowest hanging fruits you can do to greatly increase the reliability of your of your AI agent. And so that's step one. It's just, like, have a validator, which is, like, the naive implementation is, like, you ask agent to do something. It does it. Or, like, it it it submits an action candidate. You intercepts the action, then you ask it, are you sure? You know? Literally, if it's just, you sure? Already, you get a bump on your evals, which is insane. It should not be the case, but it is the case. Now if you increase if if you change this to, like, an actual prompt, which in our case now is, like, 10,000 tokens, it's a really big validator prompt. If you change that, then you you actually you're you're giving it a checklist, and that's got returning tokens. And and now it it it's way better. Like, you can get to perform above office level in our in our experience. Now then you can get can go one step further. You can you can have, like, a a federated suite of of validators, and some of these validators may be deterministic. So for example, one thing we and and the rest of the industry has found is that Sunet I mean, cloud models this year have been getting dates wrong by one day. Right? That's one way in which they'll spike it. Right? It's like, hey. You've got your AI organization. Did it gets dates wrong. It's a problem when, like us, you're building an AI executive assistant, which half its job is to schedule meetings. Okay? You can't get dates wrong. So what we did is that we we created it's it's a modular architecture we have. It's really simple. It's just like a bunch of videos, and then it's like a promise.org with a timeout for for people who know what this means. And and so everybody that there has, like, a second to, like, decide what to do. And and and and and if they've not submitted their verdict, then time's out, and the agent just submits the action. And and and we have a validator, which job it is to detect dates that were submitted in the action and to and and we we've prompted the agent to to include a weekday in the date. So it never says never says July 28. It says Tuesday, July 28. Okay? s to detect dates that were submitted in the action and to and and we we've prompted the agent to to include a weekday in the date. So it never says never says July 28. It says Tuesday, July 28. Okay? If you do that, then you can have a deterministic value detail that checks whether the date of the week matches the the day the day that was submitted, and that's just a regular expression. Instant, pre, no AI in the loop, and that's that's one of those federated validator. But even even the validators that are AI powered, you don't want initially, what we did is we were like, what if we had Sunnet for the main agent and we had, like, DeepSeek Flash for the the sub agent, the validator? But, actually, because when you when you have a cache hit, it's 10x cheaper. You you you you actually unless the model you're going to use is more than 10x cheaper, which it may in the case of DeepSeek flash, it may not because the caching is inferior. You know? If that's the case, you actually do wanna keep the same model. And so that is so much smarter. Like, it it's actually worth it to to just, like, keep the same model. That that also holds if you fork your your agent. Like, that way, can you can recycle the cache. By the way, interesting note. How do you reuse the cache when you have this validator? You know? Because the problem is that changing the toolset invalidates the cache. Okay? So the way we've done it is that the validator as as the agent has the validator action always included in its toolset, and it's unable to invoke it. We tell it, don't don't invoke this guy unless you're the validator. And if he tries, we just we just like, no. I'm not gonna I'm not gonna listen to you. I'm not the validator. And then the validator now inherits all the actions of of the agent, And it's like, you are now the validator. You may invoke this action, which you have known about the whole time. So that's how you don't break the cache with for this kind of of pattern. Yeah. And so forked forked agent's the same. You just use the same model. You know, I I've heard friends, funders who've told me I don't even wanna check if it's true. I'm sure it is, which disgusts me. If you if you take an agent and you change the model at every turn between roughly equivalent models so one turn is SunNet, one turn is Grok point five or whatever the latest is, and one tail on its, like, GPT 5.6, and one tail, it actually increases the performance. Is that ensemble? And it's literally just like a random sequence. Okay? That ensemble, for some freaking reason, outperforms any given model. I don't wanna know why, but, apparently, it works. Again, we've not even tried it because we don't wanna break the cash anyway.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.