High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Sarah Saab: belief

18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)

“You know, humans live in society and they tend to share their cultural beliefs with their tribes. And I think that's why being able to stratify the data you gather for evaluation from people, I think is quite powerful actually.”

— Sarah Saab

Source trail

Everything needed to verify it.

Speaker
Sarah Saab
Attribution
Verified speaker
Claim type
belief
Recorded
18 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…Let me paint a bit of a broader picture. So there's sort of 2 camps, if you will, in the or well, when ICE first started in this, it was all about machine learning. Now it's AI. But there's 2 camps. On the 1 hand, data is really important part of the equation, and then there's the algorithms, if you will. Most people focus on the algorithm. Most people agree that the data part is maybe perhaps the less sexy part. That's what we focus on. That's what we're all about. And it's there's a lot of agreement coming out of lots of recent papers, including constitutional AI paper, agreement that quality trumps quantity. Yet most of the models these days have been trained on an enormous corpus of data with tons and tons of noise in it, right? And even things like RLHF is just a comparative a comparison between 2 outputs with almost no reason as to why that is. So how do we bring more quality into it? In general, even your traditional machine learning model, a supervised model for a narrow target or something, back then it was irrelevant who was reviewing whether something is a cat or a dog, for example, right? Anyone could do that and you can trust that quality of that data to some degree, right? But now we're in a world where foundational models have massive amounts of capabilities and we really need to question ourselves who is producing the data that we're training on, but also the data that we're evaluating on. We're making very, very far reaching decisions to evaluate whether a model is safe, whether a model is doing well on something, or whether it converges well in in the training step, and this is all based ultimately on something that most people consider ground truth. And then but what if that ground truth is inherently noisy? What if that ground truth is inherently susceptible to tons of variants on because it's not the right people who have reviewed it, there's some bias in the order something has been reviewed, or even the interface of something is being reviewed in, or there is so many little factors that influence the quality of these labels. And, yeah, it's actually kind of fun to sort of try to unpick all of all of these different factors that go into it. There's also a really interesting piece of work being done by a group called the Collective Intelligence Project. And this group is asking these groups of people from around the world their views on a variety of AI ethics and AI safety and responsible AI topics. If you track the societal groupings, it seems you find a nice carving point for norms. You know, humans live in society and they tend to share their cultural beliefs with their tribes. And I think that's why being able to stratify the data you gather for evaluation from people, I think is quite powerful actually. Then you have these durable strata that travel through time. If we talk about the Apollo research maturity curve, I thought that was really interesting and really sort of brings to the forefront that actually this is pretty high stakes stuff. The analogy was drawn to aircraft safety Yes.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence