High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Edwin Chen: preference

7 Dec 2025 Lenny's Podcast The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI)

“Yes, so the way we really care about measuring model progress is by running all these human evaluations. So for example, what we do is, yeah, we will take Gore human annotators, and we'll ask them, "Okay, go have a conversational model.”

— Edwin Chen

Source trail

Everything needed to verify it.

Speaker
Edwin Chen
Attribution
Verified speaker
Claim type
preference
Recorded
7 Dec 2025
Publisher
Lenny's Podcast

Transcript context

…Knowing that with that in mind, how do you get a sense of if we're heading towards AGI, how do you measure progress? Yes, so the way we really care about measuring model progress is by running all these human evaluations. So for example, what we do is, yeah, we will take Gore human annotators, and we'll ask them, "Okay, go have a conversational model. " And maybe you're having this conversation with the model across all of these different topics. So you are a Nobel Prize winning physicist. So you go have a conversation about pushing different tier of your own research. You are a teacher and you're trying to create lesson plans for your students, so go talk to the model about these things. Or you're a coder and you're working at one of these big tech companies, and you have these problems every day, so go talk to the model and see how much it helps you. And because or searchers or annotators, they are experts at the top of their fields, and they are not just giving your responses, they're actually working through the responses deeply themselves, they are... Yeah, they're going to evaluate the code that it write. They're going to double check the physics equations that it writes. They're going to evaluate the models in a very deep way, so they're going to pay attention to accuracy and instruction following, all these things that casual users don't when you suddenly get a popup on your ChatGPT response asking you to compare these two different responses. People like that, they're not evaluating models deeply, they're just vibing and picking whatever response looks flashiest or [inaudible 00:21:38] are looking closely at responses and evaluating them for all of these different dimensions, and so I think that's a much better approach than these benchmarks or these random outline AV tests. Again, I love just how central humans continue to be in all this work that we're not totally done yet. Is there going to be a point where we don't need these people anymore, that AI is so smart that, "Okay, we're good. We got everything out of your heads"?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence