High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Josep M. Pujol: preference

15 Jun 2024 The Cognitive Revolution Building Brave: Private Search, One AI Layer at a Time with Josep M. Pujol

“Even though it's privacy preserving by design, still we do it on opt in because it's something that is, like, people have to be aware of.”

— Josep M. Pujol

Source trail

Everything needed to verify it.

Speaker
Josep M. Pujol
Attribution
Verified speaker
Claim type
preference
Recorded
15 Jun 2024
Publisher
The Cognitive Revolution

Transcript context

…Yeah. That's really interesting. So if you're in a country where, you have reason to be concerned about this sort of stuff, then this becomes a an extremely valuable option. Yeah. First, it's it's an opt in. Even though it's privacy preserving by design, still we do it on opt in because it's something that is, like, people have to be aware of. But they could enable it, we would not be able to it's something that's basically technically, it's called, like, record unlinkable unlinkable. Right? So any of the element that we receive, we have no way to know if those 2 elements come in the same person or not. That means that we only have individual data elements. Because even if the individual data elements, if they if we had a way to link them, we could actually create a profile. And based on that profile, it can always be de anonymized. Right? Because they will always be 1 of the data elements that somehow probabilistically or or optimistically can link to you. And then, basically, the whole session is compromised. The whole point is that we have no technical means to actually build a session so that, you know, like, there is, like, yes, somebody visited this page, but we couldn't know what happened afterwards. But, of course, that data is much less useful than profile data. Because profile data it's not that profile data is is evil. Let's put names on it just for the fun of it. Right? It's not that the Google engineer is evil, and say, I want to collect data. I'm not. But, basically, what you want is that once you collect data, you want to be reusable. You want the same data that you collect to be able to count how popular the page is, but also, like, how big how pages relate to each other or, like, how popular a particular restaurant is at what time. And you want the same dataset to be able to answer all those questions. And that's very powerful, but that's very problematic. Right? Because it can answer all these legitimate questions, but it can also answer illegitimate ones. Like, give me the search history of whoever was here at this particular time. So what we cannot do is answer these generic questions. Right? So for every question that we have, we collect 1 particular data element that only answers this particular question. So that's the difference. Like, the how is that is important. It is not so much about data or no data. It's like the purpose of the data. So how does this compare or maybe it's maybe you combine. I guess you you have a crawler too. Right? So there's Yeah. It's not that the index is entirely built from the page visits, but more that it is…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence