High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Gordon: evaluation

20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)

“Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day.”

— Andrew Gordon

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon
Attribution
Verified speaker
Claim type
evaluation
Recorded
20 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…Formula 1 cars are the absolute pinnacle of engineering. Right? Everything is perfect. Huge top speed. If you used it as your daily commuting car, you'd have an absolute nightmare. Right? And I think the same can be said for these models. Right? A model that is incredibly good on humanities last exam or MMLU might be absolute nightmare to use day to day. Most reporting on benchmarks is done on technical benchmarks these days. Right? That is where you get the model. You you give it a set of evaluations maybe on 1 theme or maybe on an exam and then you get a score and humans aren't really involved in that loop. Hey. My name is Andrew Gordon. I'm a staff researcher in behavioral science at Prolific. So I work on the sciences team, tackling questions related to humans in research, specifically online research. My name is Nora Petrova. I'm an AI researcher at Prolific. I'm tackling questions around how do we include humans, development of AI models, and evaluation AI models? How do we align them to human values, and how do we fully understand what they're capable of?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence