High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Lenny's Podcast / episode intelligence

Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar (creators of the #1 eval course)

25 Sept 2025 30 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 12evaluation 9recommendation 6uncertainty 2commitment 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

30 published records

01 / recommendation

Recommends Pachinko.

“I like to recommend a fiction book because life is about more than evals. Recently, I read Pachinko by Min Jin Lee. A really great book. Then, I also am currently reading Apple in China, which the name of the author is slipping my mind, but this is more of an exposition, written by a journalist on how Apple did a lot of manufacturing processes in Asia over the last couple, several decades. Very eye-opening.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

04 / belief

I think what's very empowering now is that product managers are doing this and can do this, and can really build very, very profitable products with this skill set.

“I think what's very empowering now is that product managers are doing this and can do this, and can really build very, very profitable products with this skill set.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

05 / belief

Just to make clear exactly what you're talking about there, one of the heads, I think maybe the head engineer of Claude Code, went on a podcast and he's like, "We don't do evals, we just vibe.

“Just to make clear exactly what you're talking about there, one of the heads, I think maybe the head engineer of Claude Code, went on a podcast and he's like, "We don't do evals, we just vibe.”
Speaker
Lenny Rachitsky
Publisher
Lenny's Podcast

06 / evaluation

If you look at the eval products, let's say the ones up until recently that some of the big labs have, they don't have error analysis. They have a suite of generic tools, cosine similarity, hallucination score, whatever, and that doesn't work.

“If you look at the eval products, let's say the ones up until recently that some of the big labs have, they don't have error analysis. They have a suite of generic tools, cosine similarity, hallucination score, whatever, and that doesn't work.”
Speaker
Hamel Husain
Publisher
Lenny's Podcast

07 / belief

The fact that your course on Maven is the number one highest grossing course in Maven, clearly there's demand and interest, and there's more people I think on your side.

“The fact that your course on Maven is the number one highest grossing course in Maven, clearly there's demand and interest, and there's more people I think on your side.”
Speaker
Lenny Rachitsky
Publisher
Lenny's Podcast

09 / belief

Now, people say the word "Evals," trying to carve out this new thing, and saying evals and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, "Hey, we need data science thinking," and AI product is helpful to have that thinking in AI products like it is in any product is my take on that.

“Now, people say the word "Evals," trying to carve out this new thing, and saying evals and then A-B testing, but if you zoom out, it's the same data science as before, and I think that's what's causing the confusion is, "Hey, we need data science thinking," and AI product is helpful to have that thinking in AI products like it is in any product is my take on that.”
Speaker
Hamel Husain
Publisher
Lenny's Podcast

11 / belief

I think you can if you want to, but the whole game here is about prioritizing. You have finite resources and finite time, you can't write an eval for everything, so prioritize the ones that are the more pesky areas.

“I think you can if you want to, but the whole game here is about prioritizing. You have finite resources and finite time, you can't write an eval for everything, so prioritize the ones that are the more pesky areas.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

12 / belief

We guarantee you, no matter what you do, if you're doing parts of these process, you're going to find ways of actionable improvement, and then you're going to iterate on your own process from there. The other tip that I would say is, we are very pro-AI.

“We guarantee you, no matter what you do, if you're doing parts of these process, you're going to find ways of actionable improvement, and then you're going to iterate on your own process from there. The other tip that I would say is, we are very pro-AI.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

14 / belief

I think people just get overwhelmed by how much time they spend up front and then thinking that they have to keep doing this all the time.

“I think people just get overwhelmed by how much time they spend up front and then thinking that they have to keep doing this all the time.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

15 / belief

" You could see the messiness of the real world in here, and the assistant just calls a tool that says transfer call, but it doesn't say anything. It just abruptly does transfer call, so it's pretty jank, I would say.

“" You could see the messiness of the real world in here, and the assistant just calls a tool that says transfer call, but it doesn't say anything. It just abruptly does transfer call, so it's pretty jank, I would say.”
Speaker
Hamel Husain
Publisher
Lenny's Podcast

17 / uncertainty

Actually, I only need to do 15." I don't know. Depends on the application and depends on how savvy you are with error analysis for sure.

“Actually, I only need to do 15." I don't know. Depends on the application and depends on how savvy you are with error analysis for sure.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

18 / belief

I think when people think A-B tests, it's like we're changing something in the product, we're going to see if this improves some metric we care about.

“I think when people think A-B tests, it's like we're changing something in the product, we're going to see if this improves some metric we care about.”
Speaker
Lenny Rachitsky
Publisher
Lenny's Podcast

19 / belief

I think that works. There's two things to that, right? One is they're standing on the shoulders of the evals that their colleagues are doing for coding.

“I think that works. There's two things to that, right? One is they're standing on the shoulders of the evals that their colleagues are doing for coding.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

20 / evaluation

Can't the AI just eval it?" That's the most common misconception, and people want that so much that people do sell it, but it doesn't work.

“Can't the AI just eval it?" That's the most common misconception, and people want that so much that people do sell it, but it doesn't work.”
Speaker
Hamel Husain
Publisher
Lenny's Podcast

21 / recommendation

If somebody is coming to you with a way to do something that's entirely new and not grounded in hundreds of years of theory and literature, then you should, I don't know, be a little bit wary of that.

“If somebody is coming to you with a way to do something that's entirely new and not grounded in hundreds of years of theory and literature, then you should, I don't know, be a little bit wary of that.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

23 / evaluation

When we're talking about LLM judge here, we're saying that this is a complex failure mode and we don't know how to evaluate in an automated way.

“When we're talking about LLM judge here, we're saying that this is a complex failure mode and we don't know how to evaluate in an automated way.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

24 / evaluation

There's a common trap that a lot of people fall into because they jump straight to the test like, "Let me write some tests," and usually that's not what you want to do.

“There's a common trap that a lot of people fall into because they jump straight to the test like, "Let me write some tests," and usually that's not what you want to do.”
Speaker
Hamel Husain
Publisher
Lenny's Podcast

25 / evaluation

I think everyone has a different mindset of evals going in, and the other thing I will say is that people have been burned by evals in the past.

“I think everyone has a different mindset of evals going in, and the other thing I will say is that people have been burned by evals in the past.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

27 / evaluation

" I'm like, "Yeah, let's look at it right now." They're surprised that I am going to go look at individual traces, and it always 100% of the time learn a lot and figure out what the problem is.

“" I'm like, "Yeah, let's look at it right now." They're surprised that I am going to go look at individual traces, and it always 100% of the time learn a lot and figure out what the problem is.”
Speaker
Hamel Husain
Publisher
Lenny's Podcast

28 / recommendation

Now, maybe you also, because these AI assistants are doing such open-ended tasks, you kind of also want to measure how good are they at very vague or ambiguous things like responding to new types of user requests or figuring out if there's new distributions of data like new users are coming and using your real estate agent that you didn't even know would use your product.

“Now, maybe you also, because these AI assistants are doing such open-ended tasks, you kind of also want to measure how good are they at very vague or ambiguous things like responding to new types of user requests or figuring out if there's new distributions of data like new users are coming and using your real estate agent that you didn't even know would use your product.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast

30 / evaluation

I think the products that are doing this, they have a very sharp sense of how well their application is performing, and people don't talk about it, because this is their moat.

“I think the products that are doing this, they have a very sharp sense of how well their application is performing, and people don't talk about it, because this is their moat.”
Speaker
Shreya Shankar
Publisher
Lenny's Podcast
Search evidence