← All source episodes Machine Learning Street Talk / episode intelligence
Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)
20 Dec 2025 8 published claims 2 attributable people
Speakers in the public record
Claim mix
evaluation 5recommendation 2prediction 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
8 published records
“Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.”
- Publisher
- Machine Learning Street Talk
“For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.”
- Publisher
- Machine Learning Street Talk
“What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field.”
- Publisher
- Machine Learning Street Talk
“I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.”
- Publisher
- Machine Learning Street Talk
“Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.”
- Publisher
- Machine Learning Street Talk
“Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you.”
- Publisher
- Machine Learning Street Talk
“Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head.”
- Publisher
- Machine Learning Street Talk
“In fact, it's not really even in the question apart from some researchers. So I mean, I I would argue that that should be just as important as how fast or smart the model is.”
- Publisher
- Machine Learning Street Talk