Evidence receipt / evaluation
Published · transcript-backedAndrew Lee: evaluation
15 May 2026 The Cognitive Revolution Three Kinds of Software Survive: Tasklet's Andrew Lee on Competing to be a Horizontal Platform
“Yeah, so caching has actually become a much bigger deal because now that the real context is in the file system, there's just a lot more tool calls that need to be done to do the basic operations of the agent because you're loading in a bunch of files and stuff.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Lee
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 15 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, okay, let's talk about that compaction, because I mean, one of the big takeaways from It might have been two conversations ago was the critical importance of caching. And you had said at the time, you know, long context is, you know, it's quite effective. Obviously it's expensive caching and especially the sort of 90% off caching pricing of Claude was like critical to enabling certain things to really work in a way that wasn't, you know, tanking your, you know, What is it? The classic meme, right? My company is dying, making it work without killing your company. So it sounds like that's changed quite a bit. Now it's much more of like a pointers, hints type of thing. How is this working? And obviously tokens are like, you know, people are spending a lot on tokens these days and I'm, you know, increasingly like hitting my limits on even the highest, you know, plan with Tasklet these days. So how are you managing context for me? What should I know? What lessons should I take away from your experience on how best to manage context in the modern moment? Yeah, so caching has actually become a much bigger deal because now that the real context is in the file system, there's just a lot more tool calls that need to be done to do the basic operations of the agent because you're loading in a bunch of files and stuff. And so We really have to make that caching work if we don't want this thing to be like crazy expensive. So that's been very much at the forefront. We came up with a new approach to context management that we shipped in December that basically works like this. You take your whole chat history, And you put it in the file system. So it's all accessible in the file system. And then you find a way to summarize that whole history in kind of a fixed length number of tokens by having recent stuff be included in like sort of high granularity, like the last thing you sent, you'll probably have most of it or all of it there. And then older things basically have like decreasing fidelity as you go back. So if you have a very long chat, the stuff in your current turn, like the current thing that's running, It's probably all going to be there, including all the thinking blocks and all the tool call responses and all the files and things are probably going to be sent to the LLM, depending on how long the run is. But for most kind of short runs, that'll be the case. The previous turn will probably mostly be there. You're going to have the full user message. You'll probably have the assistant response. You'll probably have the tool call arguments. You'll probably have the tool call responses. You'll probably have the thinking blocks. But as you go farther back, we start stripping the thinking blocks. We start stripping the tool call responses, or at least truncating the tool call responses and then stripping them. We start truncating and then stripping the tool call arguments. Then we start collapsing tool calls, and then we start shrinking down the assistant messages And then finally, we get to some LLM-based summarization. And we do this in buckets, moving back so that we can have sort of a minimal impact on caching. Basically, you want to avoid messing with prefixes as much as you can. So as you go back, basically, you get into these buckets where we have different levels of compression. And those buckets, as they get older, tend to get added to very slowly. And then once they hit a certain threshold, we shrink them down. And this system has actually worked basically pretty well. And the core thesis basically is like, you generally care a lot more about recent stuff and you trust the agent to go and like look things up when it needs to. And I would say it's not perfect. We do definitely do have people say that agents forget things. It definitely does still It costs us a lot of money to run, but I think it's generally worked. Our plan is to double down on this type of architecture. ly do have people say that agents forget things. It definitely does still It costs us a lot of money to run, but I think it's generally worked. Our plan is to double down on this type of architecture. We have lots of ideas how to improve this, but the basic approach of this decreasing fidelity as you go back and these bucketed cache-aware chunks, I think is the right approach.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.