High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Cat Wu: evaluation

23 Apr 2026 Lenny's Podcast How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)

“is important for helping the team quantify what the goal is and what their progress towards it is and what they're missing. And so I think evals is this like underappreciated thing that more PMs, more engineers should be working on.”

— Cat Wu

Source trail

Everything needed to verify it.

Speaker
Cat Wu
Attribution
Verified speaker
Claim type
evaluation
Recorded
23 Apr 2026
Publisher
Lenny's Podcast

Transcript context

…And how do you build that skill? Is it just using each, like basically understanding the limits of each model? Having like, you talked about taste, understanding, having taste into what the model maybe is capable of, what it's great and not great at, where it's changed. I think it's spending a ton of time talking and using the model. One of the things I really like to do is to ask the model to introspect on its own behaviors. So sometimes when I notice that the model does something unexpected, like for example, there's like situations where the model will make a front end change and run tests, but not actually use the UI. It's actually pretty useful to ask the model to reflect on why it did this. And sometimes they'll say that, hey, there was like something confusing in the system prompt, or I didn't realize that the front end verification was like part of this task, or hey, I delegated the verification to this subagent and the subagent didn't do the test and I didn't check its work. A lot of times just like being very curious about why the model made the decision that it did will show you what misled it so that you can fix the harness in order to close this gap. The other thing that helps is to figure out who are the users who you trust the most to give you accurate feedback about the model. Usually there's like a handful of people who are much better than others at articulating what makes a specific model or model harness combination good. There's a lot of people who will give you feedback, but not everyone's feedback is as qualified. And so finding a group of those like five people you trust is really important for getting very fast feedback. I think the third thing that is useful, but not everyone loves doing is building evals. You don't need to build hundreds of evals for them to be useful. Just building 10 great evals. is important for helping the team quantify what the goal is and what their progress towards it is and what they're missing. And so I think evals is this like underappreciated thing that more PMs, more engineers should be working on. We've covered evals a bunch. There's this trend of just like, that is the future of product management is writing evals because essentially it's what does success look like? Okay, cool. Let me actually concretely define it and then we'll know. How much of your time are you spending writing evals, would you say?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence