Evidence receipt / prediction
Published · transcript-backedAndrew Gordon: prediction
20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)
“What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field.”
Source trail
Everything needed to verify it.
- Speaker
- Andrew Gordon
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 20 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…What if we could have a a fairer approach where we actually diversely sampled and stratified folks based on how old they are and where they live and what their values are. What would be a fairer approach to understand the behavior of models and to do evaluation? So this is what Andrew and Nora have been doing at prolific. How do we know whether these models are actually good for the humans that are using them? How can we make evaluation metrics which are fairer? This is Andrew and Nora. But the problem is at the moment, I mean the field of kind of evaluations and benchmarking these models is incredibly nascent. Right? It's only been around as long as LLMs have been around in the last couple of years. And because of that, it's a kind of like a fractured field. Right? There's no there's no standard playing field for how these labs report data on on on benchmarking. You know, some may emphasize that with Grok 4 recently, we saw huge amount of emphasis on humanity's last last exam. Right? And and less so on other benchmarks. And some models will come out without any benchmarking data at all, right? And and without some kind of essentially, there's a lot of heterogeneity in how these labs report results. What it what it leads to for me is a situation where I think we're at risk of struggling to actually compare the models on any even playing field. But there's of course bigger questions as well about, you know, models are often lauded for doing the highest score in humanity's last exam. So we know from a technical perspective, the model has advanced about its rivals. But for me, and I I guess my kind of core argument in this space is if you just rely on those technical metrics, you miss half the point. Right? Like these models are designed for humans to use. At the end of the day, most of the users are humans. And simple performance on these exams doesn't necessarily correlate to a good user experience. All the frontier labs really need to start kind of having a human preference lead towards a little more front of mind alongside all the technical metrics. People are increasingly using these models for very sensitive topics and questions for mental health, for, how should they should navigate problems in their lives. And there is no oversight on that. And in any other area where these topics are discussed, there is a lot of regulation and and a lot of kind of ethical conduct built into it. Whereas here is kind of the Wild West at the moment, and some companies are taking it more seriously than others and trying to study the ways in which humans are, using the models for for more personal topics and and problems. And we've seen some pretty starky examples recently with with Groktri and Mecha Hitler. And it does raise questions about how how thin of a veneer is the safety training on top of some of these models.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.