Evidence receipt / evaluation
Published · transcript-backedSpeaker unverified: evaluation
1 Sept 2026 · 1:06:33 The Cognitive Revolution Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance
“I would worry about any memory architecture that instead of relying on lower cost embedders and rerankers is relying on multiple passes of the LLM to help you categorize and shrink the the the corpus of data that you might want to then put into the context window for the larger sort of more functional LLN call because then that's just adding up tokens as well.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- evaluation
- Recorded
- 1 Sept 2026 · 1:06:33
- Publisher
- The Cognitive Revolution
Transcript context
…write it into the memory system. And when it writes it into the memory system, the more sophisticated ones also have this notion of arbback so that if you and I are in the same job type that we get to share the memories so that if you come up with a really good memory and then I get to reuse it later, then that ends up being beneficial for both of us. So that's the second responsibility you have. But the hardest part is what you're talking about here. The way that one co-orker just wrote this to me yesterday. I want to quote him. Write, change, recall, forget. Because these things have a a half-life to them, right? Like things that are more recent are more important than ones that that took place weeks if not months ago. Maybe you can tell me about any like specific tricks or advantages for these sort of graph structures that I seem to keep making. I have people and the organizations they work for and the ideas that that I sort of associate them with. And these things all kind of point to each other, but they do it in a pretty loose way where I'm I'm right now kind of trusting the agent to hopefully notice those pointers and to the degree I'm like running maintenance, hopefully follow those pointers and do the necessary maintenance. I don't have a lot of guarantees and I suspect that there are better technologies that I could be building on that would give me a lot more robustness. Well, there's there's a couple different ways to tackle this. Like I said, the forget part is the hard part of it. But what we see our customers doing is this is where like the retrieval quality of those memory types, that's where that that non-commoditized embedding model can make a difference and where a re-ranker can make a difference. You you can get depending upon the use case, you can get like a 5 to 10% boost at your retrieval quality just by using a re-ranker on top of whatever your embedding model of choice is. Um, and that's true of architecturally. That's why we made it so easy to implement the reranker on top of the voyage embedding models when you're doing that vector search. But there are use cases where maybe you have so much data that you might have to take a hybrid approach where maybe you use maybe use a graph structure for I don't know two to six levels of data and then you get down to a leaf and you do a vector retrieval inside that leaf. We see people doing that as well. There's all kinds of information you can find on our website about how you can because we're JSON based. You can use MongoDB to build uh graph structures into your data so that you only have to go one place for a graph database, a core database, a vector database, the embedding, reranking all in one place. So you don't have to try to stitch together multiple tools yourself and have to maintain that over time. But we do see people doing that with graph structures. It's typically not as deep as like you would traditionally think of as a as a graph database need, but like I said, it's pretty common to go, I don't know, half a dozen or maybe maybe a dozen layers that maybe for bigger corpuses of data that are more heavily categorized like some of our retail customers do that d, it's pretty common to go, I don't know, half a dozen or maybe maybe a dozen layers that maybe for bigger corpuses of data that are more heavily categorized like some of our retail customers do that with product databases. If you think about how a hierarchy of products might appear on a website or on a mobile app, they might segment that first by product category using a using more of a graph style and then once you get to a particular product category, then do some vector searching within that individual category node. Are there any applications that you would point to as just being great examples of memory well implemented? Well, our biggest customer right now, they haven't been very public about how they did it. There's a company called 11 Labs out there. They had multiple agents per customer. This sort of micro agent approach where you get multiple smaller agents at the disposal for individual customers and given the number of customers they have. So they 11 Labs started life as a model provider that was doing way sophisticated speech to text and text to speech and they built a platform on top of that that is more of like a an audio editing suite that they do for that. And as part of that they they have a bunch of agents doing all kinds of editing and transformation kinds of things for their customers. And if you think about the kind of context memory that they have there an important way to to do that as well. I would worry about any memory architecture that instead of relying on lower cost embedders and rerankers is relying on multiple passes of the LLM to help you categorize and shrink the the the corpus of data that you might want to then put into the context window for the larger sort of more functional LLN call because then that's just adding up tokens as well. I mean that why things like embedders and re-rinkers have a lower token cost is again right tool for the right problem. >> Yeah, that adds a lot of latency too at this point. Even the small models can reason for quite some time before you actually get your answer from them. >> They can and so all those things add up over time and it's not just tokens but you're you're right the overall latency for the decision loop that you're into makes a big difference. I think you also have some interesting takes on build verse buy analysis. I won't even try to summarize it. Just give me your your hot takes on how people should be thinking about building verse buying. What is mature enough in the AI realm to buy and what do you even if you kind of buy part of it, what do you still have to expect that you're going to end up building or customizing enough that it kind of feels like building? Yeah. So, I have the good fortune in my job. I've been to seven countries this year to talk to probably a hundred different customers about where they are in their AI journey and they tend to fall into sort of three camps. Camp one is I bought a license for this one tool and I'm done, right? Like my AI strategy is done. I bought one thing. And while that can be a good starting point, it typically is not specific enough to help solve the problems of your overall enterprise. 'm done, right? Like my AI strategy is done. I bought one thing. And while that can be a good starting point, it typically is not specific enough to help solve the problems of your overall enterprise. Then you have a group of people who have done some PC's and maybe a couple of production deployments with varying degrees of success on ROI that have struggled with ROI. I always argue probably picked the wrong problem. Picking the right problem is really important in this space. And the way that I encourage customers to think about this is like what are the top 10 to 15 problems that you have going on in your business right now? And of those, what do you have good data for? I mean, you made a comment earlier in this conversation like data quality is a big deal. Things like bad data quality and bad security posture don't get solved by AI, they get amplified by AI. So, what's your what's your biggest problems? What do you have good data for? And then what do you already have metrics for around the problem? And that's the biggest difference that people tend to skip over when it comes to problem selection. Because if you don't already have metrics for how something is performing, you won't know if it got better. So call center use cases are very popular lowhanging fruit in enterprises because I already know what cost per call is. I already know what call volume is based on how I'm already bonusing people in those jobs. So if I introduce AI into their workflow and I see those numbers jump, I can attribute that change to the AI and I can do a back of the envelope ROI. We see the same thing with software delivery life cycle. Earlier in the year, we saw all kinds of bragging about well now because I've got clawed code or I've got or I got codeex now I can produce five times as much code as I did before. Anybody who's been doing it any length of time knows that lines of code is a terrible metric to judge the productivity of a set of developers off of. Instead, how how fast are you getting from idea to production deployment? That's the metric that matters there, not not necessarily lines of code. So, the metrics there matter and that's the difference between the folks that are in that middle state, either stuck in sort of PC purgatory and haven't quite got to production deployments. The biggest difference isn't in how they're applying the tech. It's what problems they chose to try to solve. And then you've got people that are sort of more advanced that are looking at some of these more sophisticated memory types and things where they're trying to optimize the systems that they have. And in some cases they're they're making purchases of larger platforms to help them do that. Other times they're doing homegrown. And like I said, we're we're like 18 24 months into this. There's there's no there's no one way to do this yet. There's there's no lampstack for agents yet in the way that we have with web development. We will get there having lived through that having lived through that life cycle. We'll eventually get there, but like there's there's no React and Angular. There's no LAMP stack for agents right now. You mentioned call…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.