High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Alessio Fanelli: evaluation

12 Jul 2024 Latent Space Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge

“I think the other thing to talk about here is whether or not humans are good at judging and evaluating these models.”

— Alessio Fanelli

Source trail

Everything needed to verify it.

Speaker
Alessio Fanelli
Attribution
Verified speaker
Claim type
evaluation
Recorded
12 Jul 2024
Publisher
Latent Space

Transcript context

…Sorry to disappoint, but like, we basically expect the scale of AI progress to go so fast that anyway, we will have to renew them. Of course, some of it will be for contamination issues, people trying to game the leaderboard, cheating and stuff. But a lot of it will just be because the benchmarks will have become just too easy, like the scale of the progress we did in just one year on those benchmarks is already huge. And I think that's also why the leaderboard was so important, because everybody wanted to climb the scores. And so those evaluations really saw a jump in the performance, like we've got curves on version one, which is archived, but still accessible, of model performance through time. And you can really see the steps which you had for each evaluation. I think the other thing to talk about here is whether or not humans are good at judging and evaluating these models. We're kind of slacking about it over today, but I would love to get your thoughts. It's like, at what point are we not the best people to test these models anymore? And how do you kind of balance the machine benchmarks, the MMLUs of the world, the LMSIS kind of human-driven rating, and then the AI judges, so to speak? I have many opinions about human evaluation. But I think that...…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence