Evidence receipt / evaluation
Published · transcript-backedAndrew Gordon: evaluation
20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)
“For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 20 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…There's been a lot of interesting research coming from Anthropic in that direction with regards to safety, with regards to alignments of the models and using constitutional AI and various approaches that they've that they've explored and also around mechanistic interpretability, just peering kind of behind the curtains of the models and understanding how an input produces a certain output, which can feature these concepts, which circuits get activated along the way and kind of tracing the thoughts essentially of these models and trying to isolate where potential problems may emerge. So work of this kind is very important and raising the confidence that these models will be able to handle novel situations in safe ways. I think the 1 the 1 other thing I'd say is given that Chatbot Arena is the number 1 and frankly pretty much the only human preference leaderboard out there for for LLMs, it's really important that we actually understand what's going on behind the scenes. So obviously, you know, Chatbot Arena is entirely open source. People go in, they put in a prompt, they get response from 2 different models, they then say which 1 is better. And that paper found that actually what was happening in the background is that some companies are getting access to a lot more private testing in the background than others. For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena. And it adds an element of bias into into the data, which is very very hard to get around. There are other issues that were called out and issues that we've seen ourselves, which we think dictates the need for a more rigorous and methodologically sound approach to doing these human preference datasets. Amongst us, think we had a pretty good idea. Beyond the criticisms in the Leaderboard Illusion paper, we think that it actually that paper didn't really touch on some of the other things we think should be cared about when you're doing human preference valuation. I think for me, there's 3 big areas I think where we've sought to improve. First of all, as you mentioned, sample. So obviously, the sample for the chatbot arena is anybody. Right? We don't know anything about them. We don't collect any demographic data. So they are just people going there anonymously, prompting the models and giving their preference data. Now, obviously, that's great. You get a huge amount of data, which is fantastic, but you know nothing about the people giving the data, which is fairly suboptimal. Then in terms of specificity, for anybody that's used chatbot arena, all you're doing is saying, I like this response more or I like this response more, right? In the real world, that kind of data is useless in a sense, right? It gives you a really nice way to make a nice leaderboard of AI models that tells the companies nothing about why that preference has been has been given. So in in our approach, we sought to mitigate that by actually splitting preference down into its constituent parts. So things like how helpful do people find models? What's the communication like? How adaptive do they find it? What do they think of the model's personality? And that when you get ratings on those kind of factors, what you actually get is an actionable set of results that say, okay, your model is struggling with trust or your model is struggling with personality. t when you get ratings on those kind of factors, what you actually get is an actionable set of results that say, okay, your model is struggling with trust or your model is struggling with personality. That's where you need to be focusing to really actually build a model that is good for real users in the real world. But there's no QA in the sense that, okay, I could go in and I could just say hello or I could say absolutely nothing or I could have a multi turn conversation and completely wonder from, how big is the sun to how long is a snake. Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you. 3 3 of those and you're out. So those are the kind of principles, I guess, we built the leaderboard around.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.