Evidence receipt / belief
Published · transcript-backedSarah Saab: belief
18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)
“I think people working at the very frontier of state of the art models, I think, do believe that, that that evals need to be rigorous and robust and humans have to be in the loop. But I think we are also at a very, very early stage of sort of break it and apologize later, where I think a big swathe of our industry doesn't yet think that that sort of human mediated evaluation is going to be important and perhaps will get in the way of innovation.”
Source trail
Everything needed to verify it.
- Speaker
- Sarah Saab
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 18 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I took some notes on on that Apollo research. So it was saying, what precisely does the evaluation measure? How large is the coverage of the evaluation? How robust in general are the results of the evals? What is the replicability and reliability of the evals? Are there any statistical guarantees? How accurate are the predictions about future systems? I mean, this is I I guess this is what you're saying. This is the kind of maturity curve that we need. So I think I think we think that, but I'm not sure the industry thinks that. And I think I think that's a really interesting thing. Some of the industry thinks that. I think people working at the very frontier of state of the art models, I think, do believe that, that that evals need to be rigorous and robust and humans have to be in the loop. But I think we are also at a very, very early stage of sort of break it and apologize later, where I think a big swathe of our industry doesn't yet think that that sort of human mediated evaluation is going to be important and perhaps will get in the way of innovation. So I think that tension is really important to resolve as well. Some of these statistical guarantees that that they made a lot of sense 5 years ago when we had quite specific models that you know, that that that operate in in a vertical domain or something like that. We now have these epically general language models that do all things to all people. And Enzo, I really like what you're saying is is spot on about we need to curate the data, know, and sorry you're talking about, you know, stratifying the data to reduce representational bias. But these models will be used for so many purposes like what does that what does that mean, right, in in a general sense. It also reminded me of Elon Musk. You know, there was a wonderful tweet saying that just imagine how the anthropic safety team felt when Elon published Hitler's trade to prod. And he was bemoaning afterwards that on v 7 of the foundation model, they've started curating the data better. They they started, you know, pulling out all of the racist data. But the thing is that there's still this fundamental subjectivity problem because you can strat it and you can curate, but but people from different cultures will be asking it different questions. And how do you overcome that ambiguity?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.