Evidence receipt / evaluation
Published · transcript-backedAlessio Fanelli: evaluation
12 Jul 2024 Latent Space Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge
“I think the other thing to talk about here is whether or not humans are good at judging and evaluating these models.”
Source trail
Everything needed to verify it.
- Speaker
- Alessio Fanelli
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 12 Jul 2024
- Publisher
- Latent Space
Transcript context
…Sorry to disappoint, but like, we basically expect the scale of AI progress to go so fast that anyway, we will have to renew them. Of course, some of it will be for contamination issues, people trying to game the leaderboard, cheating and stuff. But a lot of it will just be because the benchmarks will have become just too easy, like the scale of the progress we did in just one year on those benchmarks is already huge. And I think that's also why the leaderboard was so important, because everybody wanted to climb the scores. And so those evaluations really saw a jump in the performance, like we've got curves on version one, which is archived, but still accessible, of model performance through time. And you can really see the steps which you had for each evaluation. I think the other thing to talk about here is whether or not humans are good at judging and evaluating these models. We're kind of slacking about it over today, but I would love to get your thoughts. It's like, at what point are we not the best people to test these models anymore? And how do you kind of balance the machine benchmarks, the MMLUs of the world, the LMSIS kind of human-driven rating, and then the AI judges, so to speak? I have many opinions about human evaluation. But I think that...…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.