High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Speaker unverified: preference

1 Sept 2026 · 47:00 The Cognitive Revolution Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance

“the podcast or the transcript of the podcast I should say diorized so it knows what I've said and what the guest has said adds up to about a gigabyte in my case and I haven't really optimized it much at all I just kind of let the agent throw it into a database of its choosing and put whatever optimizations on it it felt like it needed to and then we did get a at one point it was like well yeah we could probably do better than keyword so we've got an embedding layer on there as well. I think full disclosure, I believe I used the Gemini embedding model for that.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
preference
Recorded
1 Sept 2026 · 47:00
Publisher
The Cognitive Revolution

Transcript context

…converted into a single vector back that represents both of those things in I guess some sort of superp position. Yes, that's exactly how it works. So you get back one array of floats exactly how you did before. It's just now you don't have to go through the iterations of figuring out the size of the chunks you need. And similarly, there's another voyage feature that all the voyage all the voyage embedding models have called retroa reasoning because the other place you have to make a decision about storage cost versus retrieval quality is with number of dimensions. So what do I mean by that? Everybody knows what two dimensions is if you've taken high school level algebra, right? XY, right? But in an embedding space, you tend to have at least 256 dimensions and sometimes as high as 2048. And the more dimensions you have, so each dimension is represented by one of those floats in that array of floats. The more dimensions you have, the richer your embedding space is and the better retrieval quality you get. But at higher dimensions, that doesn't come for free. you've got a storing 256 floats takes up less space on disk and in the index and memory than it would for 2048. So you again have to go through this iteration of what's the right number of dimensions for my use case given my what my storage costs might be. All the voyage models have a feature in them called betroka reasoning. It comes from the Russian nesting dolls. If you think about how Russian nesting dolls work, right? You've got you've got one doll and you open it up and there's another one exactly the same but smaller a smaller fidelity inside and then you keep doing that and over and over. So what what the how the how the voyage models work is when we generate let's suppose you did at 1024 let's suppose you wanted 1024 dimensions you run some tests now you want to try 512 with a traditional model you have to run your entire corpus of data through a second time at 512 but with voyage models you don't have to do that when you've run it through at 1024 vectors those floats they're ordered So if you want to try 512, you just lop off the last 512 and that's now you immediately can begin testing on the remaining 512. So again, it doesn't completely solve the problem of figuring out what the right combination of retrieval quality and storage space is, but it helps you get to the answer faster. So these are there when you take some of these things and sum there we've got three or four of these kind of features that make it easier to use and help you get to your final answer more quickly and the idea there is to give you time back in your day that you can work on your business logic instead of figuring out the plumbing. >> Yeah, that's cool. I'm a huge fan of matura anything. Uh when does this stuff become necessary? So for me, I'm basically a business of one and I try to be an early adopter of everything that I can and I do have a pretty well working I call it deep context but basically a retrieval system that allows my agent to go into kind of all of my history from the last five years essentially which is emails and Slack messages and everything I publish online and DMs across all a retrieval system that allows my agent to go into kind of all of my history from the last five years essentially which is emails and Slack messages and everything I publish online and DMs across all kinds of channels. the podcast or the transcript of the podcast I should say diorized so it knows what I've said and what the guest has said adds up to about a gigabyte in my case and I haven't really optimized it much at all I just kind of let the agent throw it into a database of its choosing and put whatever optimizations on it it felt like it needed to and then we did get a at one point it was like well yeah we could probably do better than keyword so we've got an embedding layer on there as well. I think full disclosure, I believe I used the Gemini embedding model for that. Uh, but it's like not very well optimized. How would I know if I'm really missing out on something? I I don't have like a huge evalu, right? It's just like I'm kind of vibing it with doesn't seem to be working well. Is it a matter of like data scale, scale of users? Is it about like cost? I want to optimize my inference cost and that's where I really need to get serious about how much data is is being returned. Like what are the what are the thresholds that people or obviously larger organizations cross where they're like okay can't really do it the let the agent choose its own adventure way anymore. We really need to get serious about some of these optimizations. Well, I'd return to the three things I mentioned in our first 20 minutes or so. It's it's speed, scale, and retrieval quality are the main three things. What most people do, and you're not going to offend me if this is what you did. I mean, most people start with Postgress and PG Vector, and then they choose they choose their embedding model with whatever cloud they're using. Like Gemini is prominent if you're going to use if you're going to be on Google, just like OpenAI's embedding model is is pretty popular over on Azure because of the historic relationship that those two companies have. But there comes a point in when you're when you're doing a demo when you're doing a PC it doesn't always show itself but there comes a point where when do milliseconds matter to your use case when does scale matter to your use case and that typically comes depending upon chunk size that typically comes at about 100,000 vectors is when is what I mean by scale and when does retrieval quality matter if you look at Voyage AI model hugging face has a benchmark out there called Rtech web that Voyage AI models are typically at the top of and we can get as much as a 14% improvement compared to some of those other embedding models that we just mentioned. So are there use cases for which a 14% difference in embedding model quality is the difference between a hallucination and a correct answer? And that's before you even start putting re-rankers on it, which is another way that you can boost retrieval quality without having to to do anything special to your data. So, like I said, it's it's speed, it's scale, and by scale, I typically mean in the neighborhood of 100,000 vectors and retrieval quality quality without having to to do anything special to your data. So, like I said, it's it's speed, it's scale, and by scale, I typically mean in the neighborhood of 100,000 vectors and retrieval quality out of your embedding model. Most people think embedding models are commoditized, and that is not true. There is a very big difference that you can get in retrieval quality based on what embedding model you choose. Anthropic does not have an embedding model in market. They recommend us. It's a great recommendation. You mentioned rerankers and and there was also this kind of earlier concept of sorting results out of the database. This brings to mind this concept of bitter lesson engineering that I think is kind of growing in prominence, which from my simple point of view is just like every so often you should probably kind of go through your stack and look at all the cluji extra things that you did to make things work and say like which of these do I no longer need because the model got smarter or the embedding model got better and So things are just kind of naturally working now or could naturally work now where in the past I had to kind of do all these artistal craft sort of things to to make sure that they worked. Are you seeing examples of that in the retrieval space broadly where things are in some ways getting easier or have you not really seen the bitter lesson apply to these pipelines, these production environments? Yeah, I think there's two places that I've seen that recently and one of them is with the reranking that we were just talking about. So in the spring we released dollar re-rank which is uh companion to the score fusion and rank fusion that we already talked about. So those two relate to doing a hybrid search. the dollar re-rank. Again, you would in a historical way if you would have to do your vector search once, send your results to a reranker to get them reordered in a way that is most optimal for a rag use case to then pop them into the context window. Just like we did with the score fusion and the rank fusion, we've now got a uh a stage where if you do dollar re-rank, you call the API once. we'll do both of them for you on the back end so that you only have to make one round trip to the server for that. So that's one place that we've seen some additional ease of use. The other one is a feature that we released recently called auto embeddings which is you tell us this is a better together story. You tell us what which collection you tell us which attribute on documents in that collection and you tell us which voyage model and how many dimensions you want. and we'll take care of the rest of it. Anytime an existing document with that attribute in it changes, we will automatically take that new, let's say it's text, put it through the embedding model, update the vector in the document, update the index in memory, it will do all that for you. If a new document shows up into that collection that has that attribute, we'll go through the same cycle for you. So again, trying to remove some of the plumbing so that you get some time cycles back as a developer. So you don't have to you don't have to go craft your own and…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence