High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Machine Learning Street Talk / episode intelligence

Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)

20 Dec 2025 8 published claims 2 attributable people

Speakers in the public record

Claim mix

evaluation 5recommendation 2prediction 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

8 published records

01 / evaluation

Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.

“Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

02 / evaluation

For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.

“For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

03 / prediction

What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field.

“What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

04 / evaluation

I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.

“I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

05 / evaluation

Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.

“Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

06 / recommendation

Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you.

“Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

07 / evaluation

Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head.

“Most importantly, it's based on information gains. So the way the way we pick the next pair that should occur in the tournaments is based on how much we will learn from these models going head to head.”
Speaker
Nora Petrova
Publisher
Machine Learning Street Talk

08 / recommendation

In fact, it's not really even in the question apart from some researchers. So I mean, I I would argue that that should be just as important as how fast or smart the model is.

“In fact, it's not really even in the question apart from some researchers. So I mean, I I would argue that that should be just as important as how fast or smart the model is.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk
Search evidence