High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Enzo Blindow: evaluation

18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)

“That's nice if we can align on the measurement that we consider success. And I think that's lacking to some degree because we have somehow inherently decided that chatbot arena for example is the measure of success, so people optimize for it.”

— Enzo Blindow

Source trail

Everything needed to verify it.

Speaker
Enzo Blindow
Attribution
Verified speaker
Claim type
evaluation
Recorded
18 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. And and that doesn't surprise me. They they feel mutually exclusive to me, which which is why I'm I'm constantly thinking, I don't know what your kind of prescription is here, but when people use your technology, is is is the idea that they would try and have a large foundation model that does all things to all people or if you think about it, there are so many different levers they can pull to tweak to tweak the models, you know, they could they could curate the the fine tuning data, they could, you know, stick a LoRa shim on there, they could, you know, tweak the RL post training, they could do dynamic system prompts and whatnot, there are so many different architectures that this could be leveled out in. What's the prescription? Well this is actually 1 of my 1 of her favorite things about the measurement space, because the measurement space inherently is unopinionated about any form of solution. So whether you tweak parameters, you change architectures, change algorithms, use different data sets, it doesn't matter. Right? It's the purest form of distributed optimization across everybody who tends to work on these types of problems. That's nice if we can align on the measurement that we consider success. And I think that's lacking to some degree because we have somehow inherently decided that chatbot arena for example is the measure of success, so people optimize for it. Then we decide that the next technical benchmark is the measure of success, people optimize for it. It's susceptible to Goodhat's law, right? But ultimately, the better we can design independent success measures and agree on it and make them freer freer, maybe not entirely free. I think that's maybe a bit too far fetched, but freer of of being able to to game them or to optimize for them. There are ways to remove it, right, the better we can build accountability because then we don't have to be opinionated on what model you use or what parameters you optimize for. It's we need to agree ultimately on the on the measure of of success. And with that Goodhart's Law thing, so when a measure becomes a target, it ceases to be a good measure, and and the measure is usually the proxy for the thing that we can't really quantify. So we we create a surrogate proxy for it. Should there should should we agree on a consensus of a few measures? Or that could be quite gameable. Right? Or or should we have some kind of individualized dynamic measure?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence