High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Cat Wu: evaluation

23 Apr 2026 Lenny's Podcast How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)

“ambiguous even coding is easier because you can verify the success whereas crafting the character requires a very strong sense of conviction and what who claude should be and i think she has like an incredible ability to not only mold the character but also to like articulate what the goals are what the character what's successful and what's not The other group of people who I really trust is just like the Cloud Code team.”

— Cat Wu

Source trail

Everything needed to verify it.

Speaker
Cat Wu
Attribution
Verified speaker
Claim type
evaluation
Recorded
23 Apr 2026
Publisher
Lenny's Podcast

Transcript context

…This point you made about people being very good at evaluating a model is so interesting. It's almost like a human eval of just like, okay, they understand where it's spiking or it's maybe lacking. Is there anyone specific that you want to shout out that's very good at this? Two people who I think are incredible at this are one, Amanda, who molds Claude's character. It's just like such a hard role because the task is so... ambiguous even coding is easier because you can verify the success whereas crafting the character requires a very strong sense of conviction and what who claude should be and i think she has like an incredible ability to not only mold the character but also to like articulate what the goals are what the character what's successful and what's not The other group of people who I really trust is just like the Cloud Code team. So we often have team lunches and whenever there's a new model we're testing, one of the fastest ways for us to get feedback is to just like at these team lunches, just like go to every single person and just be like, hey, what is your vibe on the model? And oftentimes we'll get feedback like, okay, this model is like not fully explaining its thinking. It's like too abrupt or like, hey, this model is like, um just like loves writing a ton of memories but like we're not sure if the memories are high quality or not or like some people will notice that okay this this model loves to test itself which is great or like this model isn't testing itself enough so that informs what data we look at to verify okay is this a larger pattern so we we have a ton of data but it is very hard to extract insights and so The feedback from this group helps us inform, okay, what are the hypotheses we want to test? And then we're able to extract data to test that. This point you made about the character of Claude, I had Ben Mann on the podcast, co-founder, and he talked about this, just like the character, the constitution of Claude is such an important part of Claude. And I didn't realize until afterwards, just like people, like with OpenClaw, actually, one of the reasons people are sad is like the personality of your Claude is like, because Claude's personality is so good and fun and and interesting unlike other models and there's and the way he put it is the personality is what makes claude so good at so many things it feels like this like trivial side thing okay it's gonna be funny and interesting and talk in a fun way but it's like so core to the success of claude is there anything good sure there about just like what people may not understand about why the character as you described and the personality is so key when you reflect on everyone you've worked…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence