High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Gordon: evaluation

20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)

“I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.”

— Andrew Gordon

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon
Attribution
Verified speaker
Claim type
evaluation
Recorded
20 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…I would just touch upon the methodology that we've used, which is trueskill. It's framework, if you will, developed by Microsoft for estimating the skill levels of players on Xbox Live. So they take into account things like randomness in games, continue changing skill levels, across time, whether someone is having kind of a fluky win streak versus a seasoned player that consistently performs well. So all of these things that we thought would be good to take into account. And, it's a very flexible system that estimates probabilities with Bayesian distributions with kind of a mean and a variance that gets narrower and narrower over time as the system kind of learns about the outcome of these battles or these comparisons. Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head. How much are how much information are they giving us? How much are they reducing the uncertainty? And we kinda order the queue of pairs according to that, and that gets us to a place of minimized uncertainty as fast as possible, as fast as we can. It's a really flexible approach. We can run separate tournaments like we've done with our demographic groups. We have around 20 demographic groups, and we've run separate tournaments for them. And we can consolidate the findings for each tournaments to obtain kind of a an overall leaderboards that is much less uncertain than any of the individual tournaments or leaderboards that we can, produce from any of the demographic groups. So it's it really allows us to slice and dice the data in any way we want, and we can easily add more demographic groups, more models over time. We're, yeah, developing in in in the open and welcoming feedback. I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models. Right? Which is all well and good. People want to play with with the state of the art. Right? Which is great. But it doesn't lead to an efficient sampling method. It it essentially means some models get a lot more battles than others. They therefore get a lot more data. They therefore get a lot better in in the arena and there's a pretty strong relationship between the number of battles and the place on the lead board. We only ever do battles based on the need from the data. So the uncertainty is high for a specific model against another specific model. We conduct a battle for that to lower that uncertainty. So it's all driven by the data. It's very computationally sound because we don't actually make any more comparisons than we need to. And it allows us to really get to a point where models are strongly differentiated based on uncertainty. If we have a certain goal with regards to uncertainty in order to fully differentiate the models at the confidence interval that we're interested in, we can just conduct more battles until we get there. The the control is in our hands. In essence, we just need to recruit more participants in order to get to that of all certainty.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence