Evidence receipt / prediction
Published · transcript-backedTim Scarfe: prediction
18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)
“You know, if we had perfect verifiers and and synthetic data generators, probably we wouldn't even need LLMs in the first place, right, because we've already solved all the problems.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 18 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Like, I'm I'm also extremely sympathetic to wanting, like, all these efforts to remove the human in the loop, like, it's costly, it's slow, it doesn't even provide, like, the best quality of data, right? There are several instances where even synthetic data might be surpassing it. But then there's instances where that's also not true, and so what we're actually working towards is a much more adaptive system. There's scenarios where human data is needed, there's scenarios where human data is very much not needed, there's scenarios where you might even have a hybrid solution, or where you meet a certain criteria, you need to have a human in the loop, where you where you almost need to define, I need this level of scrutiny now, therefore, I need higher quality input, and therefore, I it's slower and accepted slower and it has higher cost, and that is fine. So it's almost like we need this routing component there. We we ourselves, we're trying to reduce the lag or reduce the time to data as much as we can and then behind the scenes to ensure that that data can be of highest quality as possible, but at the same time, it's almost like there's a constant trade off between the quality, cost, and time. And if you want lower quality really fast at low cost, you can go with something off the shelf synthetically. If you need something really high quality, it will be the default slower and more expensive. You can get the best experts in the world to give opinions on this, right, or give input on this. But there's also an entire spectrum in between. And so how do we how do we solve for that, make the lag smaller, and make it as adaptive as possible? If I may put my cards on the table, I think the underlying problem here is that these machines don't really understand anything. That's why it's so important to get to get humans. You know, if we had perfect verifiers and and synthetic data generators, probably we wouldn't even need LLMs in the first place, right, because we've already solved all the problems. So we're we're in this we're in this intermediate phase where where we can do, you know, some of all of these different constituent parts. It sort of it seems like on the surface that we're removing more more humans from the process, but it's not entirely true. For example, reinforcement learning is a really, really good example of this, where, for example, web agents, where are now largely trained or in these environments that are often synthetically created, but there's programs that need to create these synthetic environments for RL agents to to explore in that need to be validated by humans. So we're seeing this interesting progression where humans are no longer directly involved in, like, the main system of interest, but sort of in this secondary, almost like a second order or third order abstraction moving outside, which actually is a welcome change because that means that we're focusing more on the right kind of tasks where humans are relevant, and we're focusing more on the right kind of high quality data. Right? It's completely ludicrous to create insane amount of data purely derived from humans. Like those days are gone, right? We don't need that anymore. Put the humans where they need it.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.