High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Andrew Gordon

Published podcast speaker

Claims
7
Episodes
1
Shows
1
Named items
0

Claim ledger

What Andrew said.

4 transcript-backed records

01 / evaluation

Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.

“Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

02 / evaluation

For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.

“For instance, before Lama 4 launched, we saw the meta release 27 models on the Arena. But of course, only 1 was actually reported in the end, which obviously undermines the integrity of Arena because the more comparisons you have for your model, the more access to prompts you have, the more data you have to refine a better model that's better at the Arena.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

03 / evaluation

I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.

“I think 1 of the things they pointed out in the leaderboard illusion paper was that actually some models assembled considerably higher than other models. And the I I I believe that the the folks behind Chatbot Arena said that's because people come to the arena to play with the latest models.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk

04 / evaluation

Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.

“Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.”
Speaker
Andrew Gordon
Publisher
Machine Learning Street Talk
Search evidence