High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Andrew Gordon: evaluation

20 Dec 2025 Machine Learning Street Talk Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)

“Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures.”

— Andrew Gordon

Source trail

Everything needed to verify it.

Speaker
Andrew Gordon
Attribution
Verified speaker
Claim type
evaluation
Recorded
20 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…If we have a certain goal with regards to uncertainty in order to fully differentiate the models at the confidence interval that we're interested in, we can just conduct more battles until we get there. The the control is in our hands. In essence, we just need to recruit more participants in order to get to that of all certainty. The way we've sampled for this study, we're obviously using our own participants from the prolific platform. But we've sampled effectively in based on the census data that we have for both The US and The UK. So long term vision of this would obviously be a more global product. But at the moment, what we do is we we stratify our our sample, I. Our participants were giving us this feedback by demographics like their their age, their ethnicity, their political alignment. And we have an awful lot of data from from censuses that tells us, you know, each country is made up of this certain proportion of these demographics, which essentially allows us to say that when we've amalgamated all these findings and we find that leading model, we can very confidently say that that model is preferred by as representative a set of the general public as we can possibly get. So hopefully in that sense, it's a lot more related to the real world preferences of people in the world rather than a very potentially skewed and biased subset that might be responding to the chatbot arena. Because we we ran our first 1 as an MVP, a proof of concept, that was a lot more about kind of proving that we can do this in a rigorous and methodologically sound way. When we actually ran that, we only ran it with 500 participants. It gave us a lot of insights about how we build humane, which is our leaderboard that we're working on at the moment. Now that leaderboard is actually running as we speak in the background. We're still having battles, so we expect to be able to have more data from that. But what I can say about the the first the first round that we did, models tended to perform across the board of the 6 models we tested which were leading models at the time. They performed a lot worse on personality metrics and background and culture metrics as opposed to things like helpfulness, communication, adaptiveness. What that really signals is that there's there's some some, I guess, more subjective aspect of these models which people are less impressed by. Potentially, just they were doing tasks that don't elicit a personality in the model or they don't elicit the model talking about background and culture. Also, model doesn't know their background and culture, so it's very hard to align with them. But the other possibility is that potentially models are just not very good at that. And that would be potentially an effect of the data they've been trained on. Because we know very little. Obviously models are trained on the entire internet. But when you train a model on the entire internet, do you get a personality that really represents what people want? Right? And and from from this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more kind of, I guess, objective measures. And obviously, lot of these models have undergone extensive fine tuning to tailor their personalities or tailor how they approach answering questions that are different across the different companies. But we've observed recently that there has been an increase in sycophancy or this kind of people pleasing behavior of models. And people generally don't seem to like it. 1 thing that this the results of this experiments or and and the leaderboard datasets will allow us to answer is what is the correlation between telltale signs of psychopancy and a a down downvotes in the personality metric that's does that, influence people's decisions on which model they prefer? We can, perform various types of post processing and analysis of the data to identify the levels of sycophancy that's observed in the datasets and to try to identify interesting relationships between the feedback that people gave, kind of more, model driven kind of LLM as a judge, oriented, analysis of of the conversations and kind of classification of various patterns and and and what the models exhibit. So it's quite quite interesting to see what we'll find.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence