High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Andrew Gordon

Published podcast speaker

Claims
7
Episodes
1
Shows
1
Named items
0

Claim ledger

What Andrew said.

7 transcript-backed records

01 / evaluation

Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.

“Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

02 / evaluation

For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.

“For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

03 / prediction

What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field.

“What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

04 / evaluation

I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.

“I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

05 / evaluation

Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.

“Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

06 / recommendation

Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you.

“Just topic wandering, which I don't think is a really good nuance for your models. So we built in to our structure where participants come in and they have multistep conversations with models, we built in QA that actually says, if if if you put lower low effort into your question or you start wondering, we're gonna penalize you.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

07 / recommendation

In fact, it's not really even in the question apart from some researchers. So I mean, I I would argue that that should be just as important as how fast or smart the model is.

“In fact, it's not really even in the question apart from some researchers. So I mean, I I would argue that that should be just as important as how fast or smart the model is.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk
Search evidence