High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Zevi Arnovitz: evaluation

18 Jan 2026 Lenny's Podcast The non-technical PM’s guide to building with Cursor | Zevi Arnovitz (Meta)

“Going back to your prompts, understanding what was not good enough, iterating on them and then seeing how AI's responses get better, I think that's probably one of the most important things and one of the things that divides between people who are okay with using AI and the people who actually know how to use it.”

— Zevi Arnovitz

Source trail

Everything needed to verify it.

Speaker
Zevi Arnovitz
Attribution
Verified speaker
Claim type
evaluation
Recorded
18 Jan 2026
Publisher
Lenny's Podcast

Transcript context

…Okay. Incredible. Let's wrap up this workflow. Is there anything else that's important in this workflow? And again, all this stuff is going to be available where people can just plug this stuff into their Cursor account and use it themselves. 100%. The one thing I'll say is that I think just like working in general with AI and even just like working on any product, doing constant postmortems is critical. So a lot of times we'll find all these kind of bugs or maybe Claude will fail to execute something correctly. And at the beginning when I started vibe coding, I would basically just keep running at it like running at the wall and until it worked. And once it worked, I was like, "All right, awesome. This works. Let's keep going." But I've learned over time that updating documentation and tooling is one of the biggest hacks for productivity. So when Claude will fail to do something or I'll see this really bad bug that shows that Claude really didn't understand something, I'll ask it, "What in your system prompt or tooling made you make this mistake?" And Claude will kind of like go introspective and think of what made it create that mistake. And then I'll say, "Okay, great. Let's update your tooling and documentation so that this mistake never occurs again." And I do this every time I'm either building an internal tool or anything. And I think this is just like working. If you end up doing a bunch of mistakes and then end up releasing the feature to users, so you're like, "All right, it's a big success." But going back and even when you've succeeded, looking and understanding what you did and what you could have done better is critical. And also using AI, this is probably one of the biggest unlocks. Going back to your prompts, understanding what was not good enough, iterating on them and then seeing how AI's responses get better, I think that's probably one of the most important things and one of the things that divides between people who are okay with using AI and the people who actually know how to use it. That is such good advice. So what I'm hearing is when the models do something dumb, make a mistake, you ask it to reflect on what the mistake it made was, and then you update the /command prompts with that knowledge so that in the future, it's not making that same mistake, and it just keeps getting better.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence