High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Josep M. Pujol: evaluation

15 Jun 2024 The Cognitive Revolution Building Brave: Private Search, One AI Layer at a Time with Josep M. Pujol

“Because they were like not because doing recommended system means that you have to have everything on a big cache, you know, a big memory map.”

— Josep M. Pujol

Source trail

Everything needed to verify it.

Speaker
Josep M. Pujol
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 Jun 2024
Publisher
The Cognitive Revolution

Transcript context

…Just on traffic. Right? Just on just on traffic, that's that's that's enough. Then, yeah, you might do a 1 or 2 levels deep of, like, calculations. In any way, that's, again, that's a very small contribution to the end goal. What the takeaway message here is not so much the history. It's a little bit more like that. Very often, like, the real innovation happens on something that is not as fashionable as an algorithm. It's more like a it's on the methodology. Right? And actually, that happens on Brave Search too. Brave Search is as good as it can get. It's because for for the we rely very heavily on query logs. Right? With a a query log is, in a way, is like a is a representation of what the page is about. That is not done on the anchor text, but on the queries of of another person. In a way, like Brave Search started being a recommender system engine. And that allows us to, like, to keep building query logs, which are empirically seen. And as you grow, you have more query logs, and then you can start to do, like, semantic queries. You have you've never seen, like, what is the age of Lady Gaga. Right? But you have seen Lady Gaga's age. And then you see a lot of it. Then if you've seen the other 1 and you know that it is good for this 1, it's gonna be good for the for the other query because semantically it's very it's very so you keep expanding the query semantically to do a semantic search on queries, not on content, but on queries. Right? And then you start to be able to use, like, learning models to generate some type of queries, not to try to index the content as a whole, but to to create what would be queries that would be answered to this page. Right? And then you do this, and then you add it, and then you start you become a little better. Then you start to index content of the page. So you start to use engrams. Right? Which is more conventional approach and more expensive than the other ones. So like the cost to benefit ratio is worse, but now it makes sense, right? Because it's the next step that you have to take. So then you keep adding and what you end up is having something that actually works. Does it make a nice history? No, because the story is boring. It's like your story is not, you just keep building and keep iterating until there's a you keep like doing 1% increments until you have something that works. But that's like the only way that things typically work again. And that's why I put so much effort in the example of the page rank, because that's like the story. And you know, like every step you keep adding and every step there are things that happen that were not possible before. For instance, back before before Brave, the, like, we like, this the same concept that we started was not possible 5 years before we start And why not? Because they were like not because doing recommended system means that you have to have everything on a big cache, you know, a big memory map. ted was not possible 5 years before we start And why not? Because they were like not because doing recommended system means that you have to have everything on a big cache, you know, a big memory map. So at a time where, like, the servers had 16 gigs of of RAM, max, to host something that has 2 terabytes, it takes a lot of machines. A lot of machines means a lot of latency is not feasible. However, at that time, at clicks, Amazon came. Amazon released a machine that had 1 terabyte of RAM. Suddenly, it was possible to create a cluster of 4 machines to actually put everything on memory, and everything would became very easy. Right? Where before 1 year before that, it would have required, like, 2,000 smaller servers unfeasible from the engineering perspective. And those kind of things happen all the time. Right? Another big hardware improvement that without the research wouldn't exist, NVMe hard drives. But without NVMe's hard drive, Brave Search wouldn't exist. It would not be cost effective. But NVMe's, I guess that you know what it is. Right? It's like hard drives that are super low latency on lookups. It's just faster. Is that the from…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence