High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Andrew Lee: belief

15 May 2026 The Cognitive Revolution Three Kinds of Software Survive: Tasklet's Andrew Lee on Competing to be a Horizontal Platform

“We have lots of ideas how to improve this, but the basic approach of this decreasing fidelity as you go back and these bucketed cache-aware chunks, I think is the right approach.”

— Andrew Lee

Source trail

Everything needed to verify it.

Speaker
Andrew Lee
Attribution
Verified speaker
Claim type
belief
Recorded
15 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, so caching has actually become a much bigger deal because now that the real context is in the file system, there's just a lot more tool calls that need to be done to do the basic operations of the agent because you're loading in a bunch of files and stuff. And so We really have to make that caching work if we don't want this thing to be like crazy expensive. So that's been very much at the forefront. We came up with a new approach to context management that we shipped in December that basically works like this. You take your whole chat history, And you put it in the file system. So it's all accessible in the file system. And then you find a way to summarize that whole history in kind of a fixed length number of tokens by having recent stuff be included in like sort of high granularity, like the last thing you sent, you'll probably have most of it or all of it there. And then older things basically have like decreasing fidelity as you go back. So if you have a very long chat, the stuff in your current turn, like the current thing that's running, It's probably all going to be there, including all the thinking blocks and all the tool call responses and all the files and things are probably going to be sent to the LLM, depending on how long the run is. But for most kind of short runs, that'll be the case. The previous turn will probably mostly be there. You're going to have the full user message. You'll probably have the assistant response. You'll probably have the tool call arguments. You'll probably have the tool call responses. You'll probably have the thinking blocks. But as you go farther back, we start stripping the thinking blocks. We start stripping the tool call responses, or at least truncating the tool call responses and then stripping them. We start truncating and then stripping the tool call arguments. Then we start collapsing tool calls, and then we start shrinking down the assistant messages And then finally, we get to some LLM-based summarization. And we do this in buckets, moving back so that we can have sort of a minimal impact on caching. Basically, you want to avoid messing with prefixes as much as you can. So as you go back, basically, you get into these buckets where we have different levels of compression. And those buckets, as they get older, tend to get added to very slowly. And then once they hit a certain threshold, we shrink them down. And this system has actually worked basically pretty well. And the core thesis basically is like, you generally care a lot more about recent stuff and you trust the agent to go and like look things up when it needs to. And I would say it's not perfect. We do definitely do have people say that agents forget things. It definitely does still It costs us a lot of money to run, but I think it's generally worked. Our plan is to double down on this type of architecture. ly do have people say that agents forget things. It definitely does still It costs us a lot of money to run, but I think it's generally worked. Our plan is to double down on this type of architecture. We have lots of ideas how to improve this, but the basic approach of this decreasing fidelity as you go back and these bucketed cache-aware chunks, I think is the right approach. How often does that get updated? If I have an agent that runs on a daily basis, do you try to keep the cache active from one day to the next, or is it every day we're going to have a fresh cache that will run through that whole session and all the interactions, but kind of tomorrow, you know, do we begin again or do we like on what frequency do we begin again?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence