High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Flo Crivello: belief

14 Aug 2026 The Cognitive Revolution Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses

“Like, I sometimes feel like we should be publishing. Because because I think we're doing stuff that's, like, seriously state of the art, and it's quite frequently that we do stuff.”

— Flo Crivello

Source trail

Everything needed to verify it.

Speaker
Flo Crivello
Attribution
Verified speaker
Claim type
belief
Recorded
14 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…I actually think it it's been the scaffold. It's the it's the context management. Like, this piece, like, context buildup, it's it's in a way inspired by Autowiki idea. And so we've had to do a lot of work around context management. Because once you talk to an AI employee, it's actually quite unlike just talking to an to an to, like, a or cloud because you actually expect it to keep a very rich representation of its past context and of its past memories. You really expect a level of consistency and coherence out of an AI employee that you don't expect out of your cloud. Like, cloud sort of like it's cute. Sometimes it plugs, like, previous memories about you in your chat, but, like, you don't really you don't really do, like, heavy work on an ongoing basis with it that, like, in the same way that you do with an employee. So managing all of that context, so both the memory agents that builds up the context, that's millions and millions and millions of tokens per user, and then the call agent that actually uses that context. And how does it retrieve the right information at runtime? Like, this has been, like, a major, major, major challenge. Frankly, I've sometimes been telling the team, like and I can't imagine Will's the only company thinking that. Like, I sometimes feel like we should be publishing. Because because I think we're doing stuff that's, like, seriously state of the art, and it's quite frequently that we do stuff. And, like, three to six months later, we see a paper come out and blow up about about that thing. There was one time when literally the paper was named what we had called the thing internally because it was it was obvious. And so I think, like, context and and memory management has been a really, really big part of the challenge. Reliability is always another part of the challenge. Right? Like, you want to make your your model work, and so we've we've worked quite a bit on on on reliability. Like, we've called it, a validator. It's basically a sort of it's an LLM as a judge that that triggers, like, multiple times during the tasks, but it's modular. So you have multiple LLMs as a judge fan out, and it's a sort of counsel, and then they they talk to each other, and they decide what to do. And so that's been one. The self improvement loop, I think, like, since g p t four, models have become so capable that now you can actually have self improvement loops. And so, you know, we're not the first ones to talk about it. Well, no exception. You know, we we Linde is now self improving. So it we we can literally see a curl of error rate go down into the right. You know, that's the direction you want to see error rate go down. It went down by, like literally, within the the first week of us putting the the self improvement to online, which was, like, two months ago or something, it went down by eight x, the error rate. I think I think this probably captures it. n the the first week of us putting the the self improvement to online, which was, like, two months ago or something, it went down by eight x, the error rate. I think I think this probably captures it. I think those has been those have been, like, the really big meaty chunks we've had to figure out.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence