High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Lee: evaluation

15 May 2026 The Cognitive Revolution Three Kinds of Software Survive: Tasklet's Andrew Lee on Competing to be a Horizontal Platform

“In the case of Anthropic, we're using five-minute caching, and so it doesn't stick around very long.”

— Andrew Lee

Source trail

Everything needed to verify it.

Speaker
Andrew Lee
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 May 2026
Publisher
The Cognitive Revolution

Transcript context

…How often does that get updated? If I have an agent that runs on a daily basis, do you try to keep the cache active from one day to the next, or is it every day we're going to have a fresh cache that will run through that whole session and all the interactions, but kind of tomorrow, you know, do we begin again or do we like on what frequency do we begin again? Well, there's two pieces there. There is when do things get updated on our side? Like when do we decide what that compressed history that we put into LLM looks like? And then what caching do we do on the LLM side? And the answer to the former is every time you do anything, it's sort of incrementally updated, including in the middle of runs. If you have a very long turn that uses a lot of tokens, it might actually start compressing inside that turn. And the reason that we persist that is actually calculating that could be really expensive. Running an LLM-based compaction of an older section that eats a lot of tokens. You don't want to do that every time the thing starts up. Every hour you have a trigger running and every time you have to compress a bunch of history, that can be very expensive. We keep all that around. On the model side, caching depends on the provider. In the case of Anthropic, we're using five-minute caching, and so it doesn't stick around very long. The assumption there basically is you're probably either in an active session or in the middle of a turn. which case five minute cache is enough, or you're probably waiting for next trigger to run. And like most people's triggers are not running like every, you know, half hour, they're running every few hours or every day. So the assumption there is it's not so common. And then different providers have different possibilities there. And for example, like OpenAI has much nicer caching primitives, for example, I'm happy to talk about those too. Yeah, okay, that's interesting. So it's basically constant maintenance of the higher level summaries that will be fed into the LLM and then pretty short kind of single burst style caching to actually reduce the cost of incremental calls within like one agent run. And it sounds like at least for Anthropic, that kind of is typically limited to like the cache. is hit for one run, but not hit across runs for the most part.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Named in this claim

Books, apps, tools, and people.

Search evidence