High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Josep M. Pujol: evaluation

15 Jun 2024 The Cognitive Revolution Building Brave: Private Search, One AI Layer at a Time with Josep M. Pujol

“Anyway, that's that's like 1 of the 1 of the things that that we use people for, to become more efficient and to reduce the noise, which is like the ultimate goal of any machine learning or AI.”

— Josep M. Pujol

Source trail

Everything needed to verify it.

Speaker
Josep M. Pujol
Attribution
Verified speaker
Claim type
evaluation
Recorded
15 Jun 2024
Publisher
The Cognitive Revolution

Transcript context

…Of course. Informed by the page visits? Yeah. I'd like to give specifics on that, but we may learn about our URL through web discovery project. Right? And we know this URL is popular, right, because it has to pass a form. So more than n people have to see that URL in order for us to receive that URL. So that means that URL does not belong to you. It's not your login URL, right, because because, like, you know, a 100 people across the world that access that URL. So that when we receive that URL, we only receive the URL. Right? We don't receive the content because the content, there is, like, no guarantees that the content does not contain private information that belongs to you. So then we receive the URL, which is public information, and then we actually go and fetch it. So that allows us to basically, that our index to be, like, small enough, but still because it's small because we do not crawl blindly. We don't do, a blind crawl of everything. What we do is that we crawl what people basically visit. And from there, I'm from trusted sites, so we actually do some additional crawling. But most of the crawling we do is not crawling, it's called fetching. Right? Because we actually fetched the content. It's not that we proactively crawl the whole web, because the whole web is full of noise. Right? There is like at least 10 different sites that are clones of GitHub. As odd as it sounds, it is. It's very like sites that are clones of GitHub. Right? So if you crawl crawl blindly, you just like increase by 10 your size only by adding noise. Anyway, that's that's like 1 of the 1 of the things that that we use people for, to become more efficient and to reduce the noise, which is like the ultimate goal of any machine learning or AI. I was just going to say the cleaning the dataset is often where the majority of the work goes.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence