High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Lenny Rachitsky: evaluation

14 Sept 2025 Lenny's Podcast The ultimate guide to AEO: How to get ChatGPT to recommend your product | Ethan Smith (Graphite)

“I didn't even think about this side of it when we started talking about this, but I think that's an important thing to note, is just this has nothing to do with the training data.”

— Lenny Rachitsky

Source trail

Everything needed to verify it.

Speaker
Lenny Rachitsky
Attribution
Verified speaker
Claim type
evaluation
Recorded
14 Sept 2025
Publisher
Lenny's Podcast

Transcript context

…I would assume that so there's the core model and then there's RAG. So the core model is I'm looking at common crawl on billions of web pages, and then I'm retraining the model. And if you ask something like, "What's the capital of California?" It predicts the next word, which is Sacramento. And that's based on the core algorithm, which is next-word prediction. Then there's RAG and RAG basically means search, retrieval-augmented generation. So I'm going to do a search and then I'm going to summarize the search. There are these two different things. And so most of what I'm describing is about the RAG piece, not the core model piece. To influence the core model is probably extremely hard and maybe you'll see the impact a year later. And it's probably something, some sort of obscure thing that nobody would want to do, like make a million pages that say, "Best product for X is brand." Which I don't think most people want to spend their time on. So I'm mostly focused on the RAG side, because that's the main thing that's controllable. And I think also the LLM is probably not going to say your product if it didn't show up anywhere on the RAG. So I think that's where most of the interesting stuff is from an optimization perspective. Cool. Yeah. I didn't even think about this side of it when we started talking about this, but I think that's an important thing to note, is just this has nothing to do with the training data. This is post-training, once the model's live, what it can do to find recent information using RAG, web search, things like that. Okay. Before we get into how to actually do this step-by-step, how to win at AEO. What are two or three things that you think are important for people to understand to be successful in this world just broadly? First thing is just recognizing that this is related to search. So it's LLM plus RAG, it's summarizing a set of search results usually. So LLM plus RAG, number one. Number two is topics. So in search, a landing page is targeting hundreds of keywords, which we talked about on the last podcast. So I'm not targeting one keyword like I was in 2007, I'm targeting 1,000 keywords, and each landing page needs to target that set of 1,000 keywords, and that's a topic. Same thing is true for Answer Engine Optimization. Each page is targeting hundreds, thousands, maybe tens of thousands of questions. And so I want to group all those questions, which then brings us into content, so how would I rank? How would I get my URL to rank? Or how are other URLs being decided whether or not they rank? Then answer all the questions. The more of the questions that I answer, the better. So in Google Search, if I have a landing page about website builders, the more that my page answers all of the subtopic, follow-up questions, the more likely I am to show up in Google Search. Same with chat, the more you answer all the questions, the better. If you don't answer a question, then you're probably not going to show up. And if you answer a follow-up question and subtopic somebody else is not answering, you're going to be more likely to show up. So topics, number two. The third is question research, so how do I know which questions people are asking? And that's actually pretty hard, because in search, Google just tells you what their ads API. They say, "This is the search volume for this keyword." There's a truth set from Google and ChatGPT is not giving us that, at least not yet. Maybe when they do ads, they'll give us more access to search volume, but there's no truth set. So how do we know the questions that people are asking? One way would just be to take all my search terms and change them into questions. So website builder, you can assume that what's the best website builder is probably a question that's probably asked proportional to the search volume for that keyword, so that's one. But then I mentioned that the tail is larger, and there's parts of the tail that don't exist in search. So how do we know what the tail looks like? And one strategy that you can use, is what are all the questions people are asking you on your sales calls, customer support on Reddit? Mine all those questions that exist somewhere else. Probably those same questions are being asked in chat, so that's another way to find questions. The last is citation optimization or offsite. So again, the LLM is summarizing RAG. So how do we show up with as many citations as possible? And you can break up the citations into different groups, my site, video, YouTube, Vimeo, UGC, Quora, Reddit. Tier-one affiliates like Dotdash, tier-two affiliates, blogs. So it's breaking up all those different citations and having specific strategies for each group.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence