High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Enzo Blindow: evaluation

18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)

“There's even, I believe, some countries where patients themselves are not allowed to receive transcripts from the testing facility on the chance of the patient misinterpreting the results.”

— Enzo Blindow

Source trail

Everything needed to verify it.

Speaker
Enzo Blindow
Attribution
Verified speaker
Claim type
evaluation
Recorded
18 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…Some of these statistical guarantees that that they made a lot of sense 5 years ago when we had quite specific models that you know, that that that operate in in a vertical domain or something like that. We now have these epically general language models that do all things to all people. And Enzo, I really like what you're saying is is spot on about we need to curate the data, know, and sorry you're talking about, you know, stratifying the data to reduce representational bias. But these models will be used for so many purposes like what does that what does that mean, right, in in a general sense. It also reminded me of Elon Musk. You know, there was a wonderful tweet saying that just imagine how the anthropic safety team felt when Elon published Hitler's trade to prod. And he was bemoaning afterwards that on v 7 of the foundation model, they've started curating the data better. They they started, you know, pulling out all of the racist data. But the thing is that there's still this fundamental subjectivity problem because you can strat it and you can curate, but but people from different cultures will be asking it different questions. And how do you overcome that ambiguity? It's a very prevalent problem. Right? Even even because humans are by nature diverse, we're all unique at the very, very end. There's also, by the way, an entire big blind spot to all of this, which because at the moment we have only been talking about humans directly interacting with models. There is a whole another side to the story which is where models are making decisions about humans where humans are influenced by them indirectly. And this is not where a human can give feedback directly or distill some form of preference. Right? So we need to capture ultimately or be able to measure the outcomes and the impact that these models have on humans. Right? There's a really good example of, for example, you can check whether an AI system, for example, can make the right medical diagnoses, right? That we can check as factually accurate. That same system to also transfer that diagnosis to someone who is affected by it, that is an entirely different language in that moment, right? And we should evaluate that differently. We can say the diagnosis was correct, but did it also have the correct impact on the patient in that moment. There's even, I believe, some countries where patients themselves are not allowed to receive transcripts from the testing facility on the chance of the patient misinterpreting the results. So we need to our evaluations and our measurement here needs to be nuanced enough to understand every part of the system individually, but the part that we ultimately care about is the 1 we solve for the end user here, the patient in the end. Did it elicit the meaningful change on them or not? That should be the ultimate goal, right? This is not something we can directly optimize for, but it's something we can ideally measure and build accountability on. An interesting thing that I pulled out of that Value Compass paper by Shen et al. Is the misalignment between what AIs think they are and what we think AIs are as people. And AIs think, or aspire to be, I'm going to use provocative language, autonomous thinkers. And research found that humans don't want that, right? So I think the reason I bring that up is the way you evaluate a helpful system is kind of, as you're saying, Tim, that sort of impossible problem of covering every test case in an infinite algorithm, which we will never do. But the way you evaluate a person or thinker, an autonomous being, we have loads of examples for in the world, right? Jury trials, right? Nobody expects that the moral behavior of a human is all predetermined when it's born, and we know exactly what right or wrong looks like in every case. We have loads of social structure for evaluating the behavior and agency of a person. And I do think that we kind of have to stop sort of equivocating between or maybe nobody's equivocating but me. I think I think we should assume we're building towards AGI or superintelligence or thinking creatures and work backwards as opposed to trying to sort of box in systems.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence