Evidence receipt / evaluation
Published · transcript-backedEnzo Blindow: evaluation
18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)
“Chatbot Arena is somewhere in the realm of in between because you're you're not actually ranking or you're not evaluating for 1 or the other. It is technically preference but you don't know quite whether the preference is because it said something wrong or whether the formatting was off or whether it didn't hit the cultural relativity or sensitivity or whether it wasn't adaptive enough, it doesn't tell you anything of the sorts.”
Source trail
Everything needed to verify it.
- Speaker
- Enzo Blindow
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 18 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…The benchmarking of Grok 4 is also really interesting on this because I'd be very happy to give them a lot of props for some of the stuff they're trying to do out in the open. Grok 4 wiped the floor on every benchmark, right? Including Humanity's last exam, usability experiments are revealing, you know, just leaving aside some of the more troubling findings, just usability experiments are revealing that it's not a model that feels really natural to use. So I think even in the best of cases, these benchmarking led approaches to evaluation seem to be failing us so far. I think if we try to describe the eval space a bit more holistically, LM Arena tackles a very specific part of it, right? We have on the 1 end of the spectrum, we have very technical eval that are effectively closed end solutions, it's effectively a benchmark with known outcomes, we can see whether these are hit or not. Ultimately it's a measure of accuracy or factual or correctness if you will completeness to some degree. And on the other hand of the spectrum, we have full subjectivity entirely down to the individual preference. Right? Chatbot Arena is somewhere in the realm of in between because you're you're not actually ranking or you're not evaluating for 1 or the other. It is technically preference but you don't know quite whether the preference is because it said something wrong or whether the formatting was off or whether it didn't hit the cultural relativity or sensitivity or whether it wasn't adaptive enough, it doesn't tell you anything of the sorts. Yeah, so we've developed a leaderboard which we've joyfully called humane, which is trying to address some of the limitations that were found with common leaderboards. It's the same principle ultimately, it's someone is able to take multi turn conversations with blind models effectively or blindly selected models. We're doing some a priori corrections, so we know most of their demographic and socioeconomic backgrounds in advance, so we're doing selection beforehand for the type of people that go into it. We're giving feedback right then and there as someone interacts with, the model. We're giving warnings when it's a low effort ask or when it's, unsafe potentially and so on just to, I guess, pre sanitize some of the inputs. We study the benchmarking based on the demographic stratification of the humans doing the evaluations. So you can see stuff emerge in the data like people of this age range think this model is better on helpfulness, but people of that age range disagree, and similarly with ethnicity and gender and other, strata.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.