Evidence receipt / evaluation
Published · transcript-backedSpeaker unverified: evaluation
1 Sept 2026 · 1:22:01 The Cognitive Revolution Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance
“The the cool thing about the vector search is I mean vector search at its core is taking some piece of data and mapping it into n space into some geographical geometric space and then really all a vector search is similarity like where what are the closest vectors to this new thing that I'm searching on and video is a great example audio is a great example unstructured data you just take a bunch of PDFs that you have sitting around in a shareepoint I mean those are good examples samples as well. So yeah, I think that there's there's some truth to that that we're it's not just how do I take the data and put it into a database, but now how do I find it and how do I find it in a way that is fast, is scalable, that's at good quality.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- evaluation
- Recorded
- 1 Sept 2026 · 1:22:01
- Publisher
- The Cognitive Revolution
Transcript context
…0 million which is I'm old enough to remember when that was serious money you know it's it's less than a tenth of a percent of MongoDB's overall market cap so I'm like or less than 1% I should say. So I'm just kind of wondering like what should we infer from these observations about value and defensibility in the software business over whatever passes for the long term in your mind. You know you could tell a story about like infrastructure is the big winner. Models get commoditized. incumbents can defend themselves against startups. Like what do you think are the right macro lessons to draw from this experience? I mean fundamentally what we've always been about has been how can we make the day in the life of a developer easier so that we can make it easier for them to build their business logic and spend less time worrying about the plumbing lower in the stack. I went through a description earlier in this conversation about how we implemented vector search and for us because we already had the JSON base. It was it was relatively straightforward for us to just okay you add an additional attribute that attribute is your array of floats. You generate an index based on that array of floats. So that was pretty straightforward for us to add to the existing product and take advantage of the things like the charting, the security and the data replication that we already had in the base product that some of the net new vector database companies out there have to catch up and struggle with a little bit. So that's why we chose to sort of do vector search the way that we did is because of the flexibility based into built into the document model from the very beginning. It was it was pretty straightforward for us to do that. But where it was not as straightforward and why we made the voyage acquisition was today you can still use whatever embedding model you want as long as it generates that array of floats. You put that array of floats in your document, you go create your index, you're off and running. But we saw an opportunity to to a that the market was seeing embeddings and rerankers as a commodity whereas we we saw Voyage really standing out and like I said before, Anthropic agrees with us. Um, and we could create these better together stories like the auto embeddings that I mentioned before. Another one I didn't mention is being able in Atlas to manage all your API keys in one place so that you don't have to go to one console for your core data, another console for your vector data, a third console for your embedding models, and a fourth console for your re-rinkers. You can get all that in one place. The management of it is easier over time. So if you look at the history of MongoDB features and really what our focus has been, we're all developers at heart and it's it's about making it easier for the developer ecosystem to learn and operate this these new techniques that we see in application architectures around agents whether it's the rag pipeline or whether it's the agentic memory. So how how do you make how do you make vector search more approachable? How do you reduce that lower lower that learning curve so that whether it's the rag pipeline or whether it's the agentic memory. So how how do you make how do you make vector search more approachable? How do you reduce that lower lower that learning curve so that more people can learn it more quickly and start to participate in this ecosystem? Another story that I've heard that I want to just get your reaction to and see if you think it's true in your experience is I was speaking to a founder of a Vector DB startup one time and I kind of made the case or put it to him that gez it seems to me like the incumbents are going to be able to add a vector aspect to what they're doing before you're going to be able to kind of replace everything that they're doing. And so seems like they're going to have a hard you're going to have a hard time like really displacing them. And his answer was, well, that may be true, but most of the data that is coming into our vector database has never been in a database before at all. Right? It's just been sitting out in some data lake or data warehouse or just kind of unstructured piles of documents. And now it's coming into a higher level of infrastructure and it's being made more valuable in a way that just wasn't happening at all before. >> Sure. >> Are you seeing something like that? And and what does that look like? I also noticed that there's the multimodal embedding model that supports video. So video would be potentially a great candidate for the kind of thing that has never been in a database before. Are you seeing this like phase shift of just like a much greater universe of data coming into than in previous eras? We are see we are seeing this broader ecosystem of data that that wasn't indexed before because it didn't lend itself to a traditional lexical search. The the cool thing about the vector search is I mean vector search at its core is taking some piece of data and mapping it into n space into some geographical geometric space and then really all a vector search is similarity like where what are the closest vectors to this new thing that I'm searching on and video is a great example audio is a great example unstructured data you just take a bunch of PDFs that you have sitting around in a shareepoint I mean those are good examples samples as well. So yeah, I think that there's there's some truth to that that we're it's not just how do I take the data and put it into a database, but now how do I find it and how do I find it in a way that is fast, is scalable, that's at good quality. And that's why the way that we implemented vector search and being able to within the same platform have the combination of the pre-filtering, the vector search and the lexical search, we feel like it gives us advantage and gives our developer community more levers to pull from than what some of the alternatives are. And like I said before, because it's it stands on the shoulders of the core product with the with the Atlas version, you can already deploy that any hyperscaler data center you want with good data resiliency with good security and if you need to to shard that data so that it doesn't leave particular geographies. We got all that for essentially for free because of how ter you want with good data resiliency with good security and if you need to to shard that data so that it doesn't leave particular geographies. We got all that for essentially for free because of how we implemented vector search on top of the core product. And yeah, we're we're seeing all kinds of different kinds of data get put into those documents in a way that we didn't before. You mentioned travels, I think, so far to seven countries this year where I'm I'm so kind of myopically focused on what's going on in San Francisco and Silicon Valley that I'm mindful that I may be missing important stories or differences in perspective that are going on around the world. I try to fill that gap at least with when it comes to China, but it's a big world out there. What has stood out to you in your travels this year in terms of differences of perspective on AI usage patterns, values, you name it, could be anything, but just kind of what do you think the US audience in our inward looking way is missing that the rest of the world is doing? I think biggest thing if I turn question on its head just a little bit there's a presumption in other countries that the US is ahead and doing things that other people are not and I found the opposite to be true. So I mean I live and work in out of Cincinnati so US is one of those countries. I've spent some time in both Amsterdam and London earlier in the year, but I just did a tour uh that included stops in Toronto, Bengaloo, Mexico City, and Sa Paulo. And the two most sophisticated customers I talked to this year were in Mexico City and Sa Paulo. And they assumed when we started their conversation that they had US competitors that were doing things that they weren't. And like I said, the opposite was true. I think we've reached a point with some of these technologies that like geographic barriers don't matter nearly as much as they did during like the web app era or or even during the cloud era because for the cloud era like if one of the hyperscalers didn't have a data center in your country yet, you were kind of out of luck. That's not true anymore. Like pretty much every country has at least one hyperscaler data center in it. and by extension access to models and vector databases and embeddings and rerankers like access for these things is far better than what I've seen with previous technology revolutions that we've seen. So I think understand why there would be an assumption that US companies would be ahead since so many of the bigger AI companies are US-based. But like I said, the two coolest things I've seen this year were in Mexico City and in Sa Paulo. So I think that those geographic barriers to being ahead in the market are starting to disappear. What do you think are the barriers? Why is that? I guess is it that the American companies are culturally too conservative to run as fast as some of these international companies are? Is it that is it like a leaprog story where they sort of had the international companies I mean had kind of less recent technology investment that they would have to get comfortable replacing or what's what's driving that surprising observation?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.