High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

1 Sept 2026 · 43:43 The Cognitive Revolution Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance

“It's just now you don't have to go through the iterations of figuring out the size of the chunks you need. And similarly, there's another voyage feature that all the voyage all the voyage embedding models have called retroa reasoning because the other place you have to make a decision about storage cost versus retrieval quality is with number of dimensions.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
1 Sept 2026 · 43:43
Publisher
The Cognitive Revolution

Transcript context

…le on this that we can potentially link through here. But the idea here is that if left to your own devices and figuring out your own iterations of chunk sizes, there comes a point in a graph where if you've got retrieval quality on your x on your y-axis and you've got chunk size on your x-axis. With a traditional embedding model, there comes a point where at lower chunk sizes, you have zero context. So the retrieval quality is low and it it builds, but then at some point it flattens and then it starts to degrade if your chunk sizes become too big. And that's why you have to go through these iterations to find the right combination. But what if you didn't? What if instead of sending that all as one big text blob? What if instead you send it as two? What if you sent the fidelity of the sentence you actually want and then as a second string, you send the size, you send the other contextual information. It's where the name contextualized chunking comes from. And we will figure out for you what the right chunk size is for that combination. and we will give you an a vector array that comes back that balances those things for you. When you flip that script, it turns out you can get better retrieval quality with a smaller chunk size, which otherwise is not possible if you're doing it the traditional way. You the way to get better chunk way to get better retrieval quality was with higher chunk size. But if you use this contextualized chunking, we released version three of that last summer. We just released four of that in the last six weeks. you can flip that script. Now, there still are use cases where you need to control the chunk size for different things, but if that's something that you wanted to not worry about or have to learn about, and the way I think about this is more developers will build AI agents in the next three years than did in the last 3 years. And in order to make that possible, we have to lower the learning curve. We have to make it easier for the developer ecosystem to learn how to do this. So if you don't want to have to learn how to go through the iterations of figuring out what the right balance of chunk size and retrieval quality is, instead if you could use contextualized chunking, we'll figure it out for you and you still get good retrieval quality out of it. So that's that's one of the benefits that Voyage offers that no other embedding model out on the market offers, that ease of use feature that gets you better retrieval quality. I definitely appreciate not having to worry about it, but I do want to learn a little more about it. It It sounds like this happens in sort of a testfree way at the level of the developer. So like I don't have to bring a bunch of evalu. >> So how at the level of kind of principles like how is it working in the background? And you said also I'm getting one vector back, right? So I pass in the sort of the chunk that I think I would really want to be able to zero in on and context and that gets converted into a single vector back that represents both of those things in I guess some sort of superp position. Yes, that's exactly how it works. So you get back one array of floats exactly how you converted into a single vector back that represents both of those things in I guess some sort of superp position. Yes, that's exactly how it works. So you get back one array of floats exactly how you did before. It's just now you don't have to go through the iterations of figuring out the size of the chunks you need. And similarly, there's another voyage feature that all the voyage all the voyage embedding models have called retroa reasoning because the other place you have to make a decision about storage cost versus retrieval quality is with number of dimensions. So what do I mean by that? Everybody knows what two dimensions is if you've taken high school level algebra, right? XY, right? But in an embedding space, you tend to have at least 256 dimensions and sometimes as high as 2048. And the more dimensions you have, so each dimension is represented by one of those floats in that array of floats. The more dimensions you have, the richer your embedding space is and the better retrieval quality you get. But at higher dimensions, that doesn't come for free. you've got a storing 256 floats takes up less space on disk and in the index and memory than it would for 2048. So you again have to go through this iteration of what's the right number of dimensions for my use case given my what my storage costs might be. All the voyage models have a feature in them called betroka reasoning. It comes from the Russian nesting dolls. If you think about how Russian nesting dolls work, right? You've got you've got one doll and you open it up and there's another one exactly the same but smaller a smaller fidelity inside and then you keep doing that and over and over. So what what the how the how the voyage models work is when we generate let's suppose you did at 1024 let's suppose you wanted 1024 dimensions you run some tests now you want to try 512 with a traditional model you have to run your entire corpus of data through a second time at 512 but with voyage models you don't have to do that when you've run it through at 1024 vectors those floats they're ordered So if you want to try 512, you just lop off the last 512 and that's now you immediately can begin testing on the remaining 512. So again, it doesn't completely solve the problem of figuring out what the right combination of retrieval quality and storage space is, but it helps you get to the answer faster. So these are there when you take some of these things and sum there we've got three or four of these kind of features that make it easier to use and help you get to your final answer more quickly and the idea there is to give you time back in your day that you can work on your business logic instead of figuring out the plumbing. >> Yeah, that's cool. I'm a huge fan of matura anything. Uh when does this stuff become necessary? So for me, I'm basically a business of one and I try to be an early adopter of everything that I can and I do have a pretty well working I call it deep context but basically a retrieval system that allows my agent to go into kind of all of my history from the last five years essentially which is emails and Slack messages and everything I publish online and DMs across all a retrieval system that allows my agent to go into kind of all of my history from the last five years essentially which is emails and Slack messages and everything I publish online and DMs across all kinds of channels. the podcast or the transcript of the podcast I should say diorized so it knows what I've said and what the guest has said adds up to about a gigabyte in my case and I haven't really optimized it much at all I just kind of let the agent throw it into a database of its choosing and put whatever optimizations on it it felt like it needed to and then we did get a at one point it was like well yeah we could probably do better than keyword so we've got an embedding layer on there as well. I think full disclosure, I believe I used the Gemini embedding model for that. Uh, but it's like not very well optimized. How would I know if I'm really missing out on something? I I don't have like a huge evalu, right? It's just like I'm kind of vibing it with doesn't seem to be working well. Is it a matter of like data scale, scale of users? Is it about like cost? I want to optimize my inference cost and that's where I really need to get serious about how much data is is being returned. Like what are the what are the thresholds that people or obviously larger organizations cross where they're like okay can't really do it the let the agent choose its own adventure way anymore. We really need to get serious about some of these optimizations. Well, I'd return to the three things I mentioned in our first 20 minutes or so. It's it's speed, scale, and retrieval quality are the main three things. What most people do, and you're not going to offend me if this is what you did. I mean, most people start with Postgress and PG Vector, and then they choose they choose their embedding model with whatever cloud they're using. Like Gemini is prominent if you're going to use if you're going to be on Google, just like OpenAI's embedding model is is pretty popular over on Azure because of the historic relationship that those two companies have. But there comes a point in when you're when you're doing a demo when you're doing a PC it doesn't always show itself but there comes a point where when do milliseconds matter to your use case when does scale matter to your use case and that typically comes depending upon chunk size that typically comes at about 100,000 vectors is when is what I mean by scale and when does retrieval quality matter if you look at Voyage AI model hugging face has a benchmark out there called Rtech web that Voyage AI models are typically at the top of and we can get as much as a 14% improvement compared to some of those other embedding models that we just mentioned. So are there use cases for which a 14% difference in embedding model quality is the difference between a hallucination and a correct answer? And that's before you even start putting re-rankers on it, which is another way that you can boost retrieval quality without having to to do anything special to your data. So, like I said, it's it's speed, it's scale, and by scale, I typically mean in the neighborhood of 100,000 vectors and retrieval quality…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence