High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Tim Scarfe: belief

18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)

“I think 25 was exactly the same and then I think the other 25% was nearly the same like 95% cosine distance on the embeddings or something like that.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
belief
Recorded
18 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…What we do on on our platform is we're trying to get as many humans into this process as possible. I know there's a lot of efforts to take human out of the loop, rightfully so. Right? It all has its time and place. But when we talk about alignment, and specifically value alignment, then we need to have at least in some capacity be able to capture the breadth of humanity in some capacity. Ella Marina, it's opt in, people can go there at any stage, interact with models, select their preference of which is better according to no reason aside from selecting 1 over the other, and there's no control for any population, so we can't really draw back I guess, a relationship of what are the factors at play here in a population. You could speculate that the kind of people who might participate in chatbot arena are biased into a very large degree. Right? It's good you can almost say that chatbot arena might be representative of how the tech world is perceiving the validity of these models or the preference of these models. I mean, on on this leaderboard illusion, that was Maazi Fadi and then Sarah Sarah Hooker and Shivalee Kasinghe at Cohere and a few other people. We did a video on that recently. It's become the de facto standard for benchmarking large language models, and it has so many problems. And it's now after this investment, it's worth $600,000,000. The insane selection bias, the the bias in sampling, know, the the the private pools, you know, where folks can kind of get more matches and and then they can take that training data and they can fine tune on the training data. And also just the foundation models, you know, from Google and Meta and and XAI and and so on, they just get given more matches. It's just incredibly unfair. And and as you were just saying before, even the folks when they put their prompts in, something like 50% of the prompts are basically carbon copies of the last month. I think 25 was exactly the same and then I think the other 25% was nearly the same like 95% cosine distance on the embeddings or something like that. So to me, that is an example of a superficially good rank but it's flawed in so many ways. The benchmarking of Grok 4 is also really interesting on this because I'd be very happy to give them a lot of props for some of the stuff they're trying to do out in the open. Grok 4 wiped the floor on every benchmark, right? Including Humanity's last exam, usability experiments are revealing, you know, just leaving aside some of the more troubling findings, just usability experiments are revealing that it's not a model that feels really natural to use. So I think even in the best of cases, these benchmarking led approaches to evaluation seem to be failing us so far.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence