Evidence receipt / prediction
Published · transcript-backedEthan Smith: prediction
14 Sept 2025 Lenny's Podcast The ultimate guide to AEO: How to get ChatGPT to recommend your product | Ethan Smith (Graphite)
“I knew that because I created spam in 2007, and I knew what Google did about it and how, and I knew the exact same thing was going to happen.”
Source trail
Everything needed to verify it.
- Speaker
- Ethan Smith
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 14 Sept 2025
- Publisher
- Lenny's Podcast
Transcript context
…Ethan, let me just say, I'm learning so much from this conversation, what a fun thing. I could see, it's just clear how much you love this stuff, and just how nerdy and deep you get into it. And it's just fun to talk to someone that's so deep and knowledgeable about all these things, so thank you for sharing all this with us. I'm going to go in a slightly different direction. There's this whole world of AI content, people generating content with AI, generating landing pages. Just like, "Oh my God, SEO is never going to just generate all this stuff. AI is going to make all this stuff easier." You guys did a really big study on how that works, whether it's a good idea to generate content with AI. Can you just talk about what you learned from that, and how people should think about AI in generating content? Yes. So I remember when ChatGPT launched and Brian Balfour posted on LinkedIn, "What do you people think that is going to happen from ChatGPT and AI?" And my immediate response is spam, so just lots and lots of spam, especially SEO spam. And then there was a whole industry around AI-generated content, and I knew immediately that it wouldn't work. And the reason why I knew it wouldn't work, and when I say AI-generated content, I mean automated content with no human-in-the-loop. So I think that the future of content is clearly AI-assisted. Clearly, you and I will be using AI to help us write, so it's not no AI at all, but it's not 100% generated with AI. I immediately knew that it wouldn't work. Why did I know that? I knew that because I created spam in 2007, and I knew what Google did about it and how, and I knew the exact same thing was going to happen. So what I did in 2007 is I and all the other shopping comparison people scraped all each other's content, reviews, chopped it up, scraped content, 100 million search pages, snippets, and it worked really well. And then it stopped working, and then all those companies disappeared. I knew that was exactly what's going to happen with AI-generated content. And so from the beginning, I've not focused on AI-generated content. Many people have, but I don't know, so maybe it does work. There's lots of case studies about it working. So let's do the study, let's do an analysis. So we took, we looked at both Google and at ChatGPT where we took thousands of searches and thousands of questions, and we put those searches into Google Search. We put those questions into chat and the ChatGPT, and then we looked at the citations or the Google Search results. Then we looked at an AI detector. So we used Surfer SEO's AI detector. Now, when I tell people this, they say, "Well, you can't detect AI." So then we evaluated the efficacy and the accuracy of the AI detector. So we did that by generating thousands of AI-generated articles and it was very predictive. And then we looked at real articles, we did that two different ways. One way is we write real articles, and the other is we took a random sample of 100,000 URLs from Common Crawl over the last five years. And then we looked at the AI detector before ChatGPT was launched, so it necessarily was content not created by a human. And then the false positive rate was around 8%, so basically the AI detector is very accurate. So we took that, then we ran it on the content. So then what we saw was around 10% to 12% of content in Google Search, and then ChatGPT or AI-generated, 90% are not. And we ran a correlation analysis showing the exact same thing. So we essentially did a very rigorous study showing that AI content does not work. AI-assisted content edited is great. We do that sometimes, other people do that, that is clearly the future of content. So that does work and should work and that's good, but purely 100% AI-generated does not work. ntent edited is great. We do that sometimes, other people do that, that is clearly the future of content. So that does work and should work and that's good, but purely 100% AI-generated does not work. So then the second thing that we did was we found that, this was unexpected, but we found that there's more AI-generated content on the internet than human-generated content. So back to the Common Crawl study, we looked at 100,000 different URLs over the past five years. And then you can see this curve where AI-generated is now higher than human-created. So there's more AI-generated content on the internet than human-generated content, which is disturbing. So then let's say that AI-generated content did work. If AI-generated content worked, then everyone would do it. Just like in 2007, shopping comparison sites, if I can scrape my content, why would I pay anyone to write it? I'll just scrape it from you and I'll chop it up. So then everyone will do that, and then it will go from most content is AI-generated to almost all of the content is AI-generated. Then what will happen if that works, is that Google now becomes a search engine for ChatGPT responses. So if Google's a search engine for ChatGPT responses, there's no reason for Google to exist. Just go to ChatGPT, which is the exact same thing that happened in 2007. Google said, "I see all these shopping comparison search engines showing up in my search results. So I'm essentially a search engine for search engines." I should be showing the TV in my results. I shouldn't be showing other vertical search engines, so I'm going to get rid of them and I'm just going to go straight to the product. The same thing will be true for ChatGPT. Now for ChatGPT, let's say that ChatGPT ranks its own derivatives in its citations, so then you have this infinite loop of derivatives. So I go to ChatGPT, I say, "Generate 10 articles." I put those articles into the citations and then I say, "Summarize these citations that were derivative." And then I keep on doing derivatives of derivatives, and then you have an infinite loop of derivatives, and now AI is summarizing itself. There's a paper about this called Model Collapse. So again, there's the core algorithm and then there's the RAG piece. So the core algorithm, a group did a study showing model collapse, which was what if you feed in AI derivatives into the model and train the core model on the derivatives? And then what happened was you had all these problems, hallucinations, things break very quickly. Okay. So then we did a study on what if you feed derivatives into the RAG piece? So generate 10 derivatives, put that into RAG, summarize that. And then generate 10 more, and then summarize my summarizations, infinite loop of derivatives. What happens? And so what happens is there's a wisdom of the crowd. The LLM is summarizing the opinion of many people. So if you ask a question like, "What's the best flavor of ice cream?" There's not one answer, there's thousands of opinions. So the LLM is summarizing these many, many opinions in this wisdom of the crowd.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.