Evidence receipt / prediction
Published · transcript-backedLenny Rachitsky: prediction
24 Aug 2025 Lenny's Podcast Inside the expert network training every frontier AI model | Garrett Lord (Handshake CEO)
“Okay, so before we get into the insane trajectory that your data labeling business is on, which is just an amazing story that I think a lot of founders and product teams that are trying to navigate this AI disruption that's happening will have a lot to learn from.”
Source trail
Everything needed to verify it.
- Speaker
- Lenny Rachitsky
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 24 Aug 2025
- Publisher
- Lenny's Podcast
Transcript context
…Yeah, thanks for having me. A long-time subscriber. I appreciate that. Okay, so before we get into the insane trajectory that your data labeling business is on, which is just an amazing story that I think a lot of founders and product teams that are trying to navigate this AI disruption that's happening will have a lot to learn from. I want to first help people understand what the hell data labeling actually is. Just like, what are people actually doing? Why is this so valuable? Some of the most, I don't know, fastest-growing companies in the world today, including you guys are just, this is what you do. Clearly there's something really important here. I sort of understand it, probably not really. I think a lot of listeners feel the same way. So let me just ask you this, what is data labeling actually? What are people actually doing? And then, just why is this so valuable to frontier AI labs? Yeah. So, I think it's helpful to take a step back of what does training a model look like? So, there's really two primary functions. There's a pre-training and a post-training process in training a model, and for a long time these AI providers, or LLMs, or Frontier Labs we're focused on basically sucking up more and more information on the pre-training side of the house. And that's basically the entire corpus of written human knowledge. So, that's not just written, but every YouTube video, every book, basically the pursuit of sucking up everything that was on the internet, and that was the pre-training side. And there was a lot of gains from pre-training, like models continue to get better. And about 18 months ago, 24 months ago, we started to really see an asymptoting of gains coming from, because they had essentially sucked up all of the knowledge on the internet. And so, labs really shifted towards most of the gains now coming from the post-training side of the house. And what post-training is, is it's augmenting and improving the data they have across every discipline or capability area that they care about. So take coding, or mathematics, or law or finance, they are focused on collecting high quality data that really improves the state of our capabilities, their models, and you can see a lot of these popular benchmarks on what are called model parts. When Llama IV is released, you'll see the benchmarks across various domains, and each one of the research teams inside of the labs have different use cases. Basically they're running experiments, almost think like the scientific process. They have a hypothesis around how to improve the model. They're trying to collect small pieces of data to see if that hypothesis works out. If that hypothesis is proving true, then they expand the overall collection of the data in that advert. And it could look like reinforcement learning environments, it could look like trajectories, it could be audio and multimodal, it can be text-based like prompt-response pairs. It can also be reinforcement learning with human feedback, which is like preference ranking data. And so, that's the state of the art of models. And most of the gains that are happening from models right now are coming from the post-training side of the house. And there's just an incredible amount of demand to stay at the absolute frontier of where models are going.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.