High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Nora Petrova: evaluation

20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)

“Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head.”

— Nora Petrova

Source trail

Everything needed to verify it.

Speaker
Nora Petrova
Attribution
Verified speaker
Claim type
evaluation
Recorded
20 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…t when you get ratings on those kind of factors, what you actually get is an actionable set of results that say, okay, your model is struggling with trust or your model is struggling with personality. That's where you need to be focusing to really actually build a model that is good for real users in the real world. But there's no QA in the sense that, okay, I could go in and I could just say hello or I could say absolutely nothing or I could have a multi turn conversation and completely wonder from, how big is the sun to how long is a snake. Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you. 3 3 of those and you're out. So those are the kind of principles, I guess, we built the leaderboard around. I would just touch upon the methodology that we've used, which is trueskill. It's framework, if you will, developed by Microsoft for estimating the skill levels of players on Xbox Live. So they take into account things like randomness in games, continue changing skill levels, across time, whether someone is having kind of a fluky win streak versus a seasoned player that consistently performs well. So all of these things that we thought would be good to take into account. And, it's a very flexible system that estimates probabilities with Bayesian distributions with kind of a mean and a variance that gets narrower and narrower over time as the system kind of learns about the outcome of these battles or these comparisons. Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head. How much are how much information are they giving us? How much are they reducing the uncertainty? And we kinda order the queue of pairs according to that, and that gets us to a place of minimized uncertainty as fast as possible, as fast as we can. It's a really flexible approach. We can run separate tournaments like we've done with our demographic groups. We have around 20 demographic groups, and we've run separate tournaments for them. And we can consolidate the findings for each tournaments to obtain kind of a an overall leaderboards that is much less uncertain than any of the individual tournaments or leaderboards that we can, produce from any of the demographic groups. So it's it really allows us to slice and dice the data in any way we want, and we can easily add more demographic groups, more models over time. We're, yeah, developing in in in the open and welcoming feedback. I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models. Right? Which is all well and good. People want to play with with the state of the art. Right? Which is great. But it doesn't lead to an efficient sampling method. It it essentially means some models get a lot more battles than others. They therefore get a lot more data. They therefore get a lot better in in the arena and there's a pretty strong relationship between the number of battles and the place on the lead board. We only ever do battles based on the need from the data. So the uncertainty is high for a specific model against another specific model. We conduct a battle for that to lower that uncertainty. So it's all driven by the data. It's very computationally sound because we don't actually make any more comparisons than we need to. And it allows us to really get to a point where models are strongly differentiated based on uncertainty.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence