Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
15 Jun 2024 The Cognitive Revolution Building Brave: Private Search, One AI Layer at a Time with Josep M. Pujol
“Going back to this plumbing example, and I think trying to generalize from that to what I think a lot of people who listen to this show are working on, I think in a lot of scenarios, there's like a proprietary dataset that is like too big for people to like page through.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 15 Jun 2024
- Publisher
- The Cognitive Revolution
Transcript context
…But this actually is I I I can tell you another story that we that might actually work because in a way, like the the level of noise that you have on vertical search is artificially reduced. I don't know if it makes sense with that. It makes sense. Yeah. Because because the problem is the the the problem is always the noise. Right? The uncertainty. So when you are doing a vertical something, vertical search, or you're working a particular vertical, noise is artificially reduced. Right? Because in a particular domain, a particular word has only 1 particular meaning, or particular token has only 1 particular meaning. And so when you work only on a vertical, it's very easy to work on. And, actually, we 1 of the things that when we first started to tackle problem search, the initial thing to to try was like, hey, my search can be split into verticals. Right? So if we are like 20 people, why don't we just do vertical? You do like weather search. You I do city search. You do famous people search, car search. Like, you come up with, like, 100 different verticals. So what ended up happening is that all the 100 verticals were perfectly. Right? But there was, like, no way to put them together in a in a way that made sense. Right? Because what we actually did intentionally is, like, we pushed the problem, the real problem. We pushed it under the rack. Right? On the module that the mixing module that will come in the future. On the mixing module is a problem. Right? Because on the movie vertical, Ted has only 1 meaning. Right? The movie Ted. Could be Ted 1, Ted 2, Ted 3. Right? But now forget about that you are on the movie vertical. Think of the whole web. But that what does it mean? The movie, name of someone, the conference, a part, you see. Right? Which 1 it is? That's like the that's why search is complicated because it's it's general purpose. And I I have the same issue. Right? The narrower the domain, the easier it is because there is less noise. Does that assume though that you are working with enough data? Going back to this plumbing example, and I think trying to generalize from that to what I think a lot of people who listen to this show are working on, I think in a lot of scenarios, there's like a proprietary dataset that is like too big for people to like page through. Right? So they do need some search to access it effectively. But then at the same time, it's not so much data that they could train the model from scratch on it. Maybe they could fine tune. And so I feel like a lot of people find themselves in that in between zone where they're like, I need to use some sort of pre trained open source model at least as a base. And those models, like in this plumbing scenario, everything is in a very narrow section of the, like, broad general purpose embeddings models. And then things like parts numbers, as you said, are are not even really semantic at all in a lot of cases. So we are finding that we do need a hybrid search if only to handle these things like That that…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.