Evidence receipt / evaluation
Published · transcript-backedDiarmuid Gill: evaluation
9 May 2026 The Cognitive Revolution Milliseconds to Match: Criteo's AdTech AI & the Future of Commerce w/ Diarmuid Gill & Liva Ralaivola
“All of that process gets done in milliseconds because we use a lot of caching, we've trained the models offline, then the inference happens at real time in really low latency.”
Source trail
Everything needed to verify it.
- Speaker
- Diarmuid Gill
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 9 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Can I dig in a little bit more on the core models that you guys are using to make predictions. And I'd love to understand the architecture of this better. I think for calibration, anybody who's listening to this feed is going to have a conversational familiarity at least with how large language models work. So we know that they're generating a token at a time. We know that the inputs get embedded and we know the kind of mechanics of the forward pass and all that stuff, right? And we know it's auto progressive, blah, blah, blah. This strikes me as a very different world. And I don't have nearly as much intuition for what the models are that are driving these things. I do know that they have to be a lot faster because the ad's gotta show up really quickly on the page. And then I know also that there's a pretty challenging matching problem in there somewhere because I've got millions of, you've got, we've got, society collectively has got millions of these profiles of individuals. And then also, as you said, into the tens of thousands of advertisers. And I don't know how much pre-computing is done or whatever, but it has to happen pretty quick on the load of a page. So could we kind of break down how big are these models? What do the inputs look like? You could imagine something very large and sort of very sparse set of inputs. But I guess it doesn't seem plausible, but it's like, here's all the websites, and here's which ones this user visited, right? That doesn't seem like it works. So there's got to be some sort of tokenization or something that is kind of bringing the user profile into a manageable state size so that it can be used as an input. I'm not even sure if I'm quite asking the right questions here. Tell me what this looks like under the hood. Yes, maybe I can take a quick stab at it and then Liva can take it down into more detail. So Liva actually referenced this earlier. So every single time that we want to, when we get an opportunity to show an ad, so that opportunity actually goes to multiple different ad tech providers who are all acting as kind of delegates on behalf of the actual advertiser themselves, whether it are brands or advertisers. And so the amount that we bid is based on how valuable that opportunity is to the advertiser. Effectively, how likely the user is to click on that ad and go back to the website and buy the product. And the way we evaluate that is we, through the mechanism we talked earlier, we see what products the users are interested in, what they've looked at, what they clicked through, what they've seen, what they buy, what they don't buy, and so on. As the display opportunity comes up, so we see the ID that we mentioned in the cookie, and then we take a look at all of the different products that that person has seen or whatever audience segments they belong to. And each one we can say, okay, based on all the different features we put into the model. So, you know, the products, the previous purchase history, the context of the website, the device they're on, a couple of other things that come in, and there's actually probably I'm not sure, it's like 150 different features we can take in. And each of those go into this calculating as part of this massive equation, which will tell us the likelihood that person is to click, the likelihood they are to click through to the website and eventually do a purchase. And all of that comes out to a value which we bid. If we win the opportunity, then we have to say, well, which products do we show and how do we do all of that kind of stuff? All of that process gets done in milliseconds because we use a lot of caching, we've trained the models offline, then the inference happens at real time in really low latency. So one thing that is important regarding all the data that we have, like the visited websites, the products that were shown, clicks, et cetera, and then the model, the question, as I said before, is kind of, let's reduce it to a classification, a binary classification problem. One of the main tasks for people who are have tried to do some machine learning, how you're going to encode and how you're going to represent the data. So I'm going to do that in two steps. The first one is going to, I'm going to talk about the legacy old models that we used to have and where we are now and that where we've been for a couple of But before, there was this question about all the products, the website that you visited, you have to encode them. And you have to encode them so that the way you encode, so the vector you're going to use to represent all those past information still carries a meaning. If you just encode them in a silly way, you're going to lose a lot of information. So before that, There was a choice before because of speed of computation and because you have crazy intuition about the type of model. It was like a spot. a representation, like a very huge vector with 2 to the 12th number of inputs with 1, 0, 1, 0, 1, 0, because we can do very fast computations on those parse vectors. But that was a way to represent the data that we used to have. And we just learned from that vector what is called the logistic regression model. So it's a linear model. You can think of just one neuron with a lot of neurons coming in if you have the learning analogy in mind. And we used to have that and we learned the model and it was very fast, even though it was sparse. There are many libraries to do like sparse matrices and sparse vectors. But then it was one thing that was very manual and built the features before. And of course, reasons why, for instance, the Crypto iLab and the head of was created was to say, okay, maybe It's not sustainable to have to craft new features and to think about how we're going to represent data each time, because the cookies can change, the information that we have, and for instance, with LLMs is going to change. So how can we proceed with more modern techniques? So it's... It was the intent and the goal of the Criteo Lab to bring deep learning. So it was created in 2018, and it was precisely the objective to say, okay, let's go to the next level, not have handcrafted features, but rather have them computed from the data. So Now, before we had like two to the 12 or to the 20s, depending on the encoding sparse vectors. Now, essentially, we have something like between 200 to the 1000 features that are automatically computed by one of the proprietary algorithm that we have, which is called DeepKNN, that computes deep learning features from which we, on top of which we do, we do learn some other models that are going to do those classification tasks.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.