Evidence receipt / recommendation
Published · transcript-backedAndrew Gordon: recommendation
20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)
“Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 20 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I think the 1 the 1 other thing I'd say is given that Chatbot Arena is the number 1 and frankly pretty much the only human preference leaderboard out there for for LLMs, it's really important that we actually understand what's going on behind the scenes. So obviously, you know, Chatbot Arena is entirely open source. People go in, they put in a prompt, they get response from 2 different models, they then say which 1 is better. And that paper found that actually what was happening in the background is that some companies are getting access to a lot more private testing in the background than others. For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena. And it adds an element of bias into into the data, which is very very hard to get around. There are other issues that were called out and issues that we've seen ourselves, which we think dictates the need for a more rigorous and methodologically sound approach to doing these human preference datasets. Amongst us, think we had a pretty good idea. Beyond the criticisms in the Leaderboard Illusion paper, we think that it actually that paper didn't really touch on some of the other things we think should be cared about when you're doing human preference valuation. I think for me, there's 3 big areas I think where we've sought to improve. First of all, as you mentioned, sample. So obviously, the sample for the chatbot arena is anybody. Right? We don't know anything about them. We don't collect any demographic data. So they are just people going there anonymously, prompting the models and giving their preference data. Now, obviously, that's great. You get a huge amount of data, which is fantastic, but you know nothing about the people giving the data, which is fairly suboptimal. Then in terms of specificity, for anybody that's used chatbot arena, all you're doing is saying, I like this response more or I like this response more, right? In the real world, that kind of data is useless in a sense, right? It gives you a really nice way to make a nice leaderboard of AI models that tells the companies nothing about why that preference has been has been given. So in in our approach, we sought to mitigate that by actually splitting preference down into its constituent parts. So things like how helpful do people find models? What's the communication like? How adaptive do they find it? What do they think of the model's personality? And that when you get ratings on those kind of factors, what you actually get is an actionable set of results that say, okay, your model is struggling with trust or your model is struggling with personality. t when you get ratings on those kind of factors, what you actually get is an actionable set of results that say, okay, your model is struggling with trust or your model is struggling with personality. That's where you need to be focusing to really actually build a model that is good for real users in the real world. But there's no QA in the sense that, okay, I could go in and I could just say hello or I could say absolutely nothing or I could have a multi turn conversation and completely wonder from, how big is the sun to how long is a snake. Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you. 3 3 of those and you're out. So those are the kind of principles, I guess, we built the leaderboard around. I would just touch upon the methodology that we've used, which is trueskill. It's framework, if you will, developed by Microsoft for estimating the skill levels of players on Xbox Live. So they take into account things like randomness in games, continue changing skill levels, across time, whether someone is having kind of a fluky win streak versus a seasoned player that consistently performs well. So all of these things that we thought would be good to take into account. And, it's a very flexible system that estimates probabilities with Bayesian distributions with kind of a mean and a variance that gets narrower and narrower over time as the system kind of learns about the outcome of these battles or these comparisons. Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head. How much are how much information are they giving us? How much are they reducing the uncertainty? And we kinda order the queue of pairs according to that, and that gets us to a place of minimized uncertainty as fast as possible, as fast as we can. It's a really flexible approach. We can run separate tournaments like we've done with our demographic groups. We have around 20 demographic groups, and we've run separate tournaments for them. And we can consolidate the findings for each tournaments to obtain kind of a an overall leaderboards that is much less uncertain than any of the individual tournaments or leaderboards that we can, produce from any of the demographic groups. So it's it really allows us to slice and dice the data in any way we want, and we can easily add more demographic groups, more models over time. We're, yeah, developing in in in the open and welcoming feedback.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.