Evidence receipt / evaluation
Published · transcript-backedEnzo Blindow: evaluation
18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)
“For example, reinforcement learning is a really, really good example of this, where, for example, web agents, where are now largely trained or in these environments that are often synthetically created, but there's programs that need to create these synthetic environments for RL agents to to explore in that need to be validated by humans. So we're seeing this interesting progression where humans are no longer directly involved in, like, the main system of interest, but sort of in this secondary, almost like a second order or third order abstraction moving outside, which actually is a welcome change because that means that we're focusing more on the right kind of tasks where humans are relevant, and we're focusing more on the right kind of high quality data.”
Source trail
Everything needed to verify it.
- Speaker
- Enzo Blindow
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 18 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…If I may put my cards on the table, I think the underlying problem here is that these machines don't really understand anything. That's why it's so important to get to get humans. You know, if we had perfect verifiers and and synthetic data generators, probably we wouldn't even need LLMs in the first place, right, because we've already solved all the problems. So we're we're in this we're in this intermediate phase where where we can do, you know, some of all of these different constituent parts. It sort of it seems like on the surface that we're removing more more humans from the process, but it's not entirely true. For example, reinforcement learning is a really, really good example of this, where, for example, web agents, where are now largely trained or in these environments that are often synthetically created, but there's programs that need to create these synthetic environments for RL agents to to explore in that need to be validated by humans. So we're seeing this interesting progression where humans are no longer directly involved in, like, the main system of interest, but sort of in this secondary, almost like a second order or third order abstraction moving outside, which actually is a welcome change because that means that we're focusing more on the right kind of tasks where humans are relevant, and we're focusing more on the right kind of high quality data. Right? It's completely ludicrous to create insane amount of data purely derived from humans. Like those days are gone, right? We don't need that anymore. Put the humans where they need it. We trafficking in human data at some point somewhere without always owning up to that. I think the most worrying version of that is let the user in production find the problems. In some cases, that's fine, and in some cases, that's very scary. And I think, Tim, if we thought they could understand, we would also hold them to account for their actions, but we can't. And since humans are being held to account for the actions of AI models, I still think the onus is on us to ensure that they're behaving the way we expect. But once they understand, we will hold them to account, right? There will be real stakes for them as people in what they do or think. But we're we're not there yet.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.