Evidence receipt / belief
Published · transcript-backedLenny Rachitsky: belief
27 Mar 2025 Lenny's Podcast How to win in the AI era: Ship a feature every week, embrace technical debt, ruthlessly cut scope, and create magic your competitors can't copy | Gaurav Misra (CEO and co-founder of Captions)
“I was actually just looking at it where you put a photo in, I think, and it just creates a video of this person talking in all these different ways.”
Source trail
Everything needed to verify it.
- Speaker
- Lenny Rachitsky
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 27 Mar 2025
- Publisher
- Lenny's Podcast
Transcript context
…just happens to produce semi-realistic outputs, but it actually stops at some point because it's not clear how it becomes generalizable in every situation. It has to be trained on people individually. So you might ingest a little bit of video of you and then you can generate you. And so it's a different technology and a different outcome, essentially. And a bunch of companies using this type of model, a bunch of companies are doing general text to video with no audio today. These are large generative models and they have the capability to do more, but that frontier just hasn't been reached yet. I think there's no doubt in anybody's mind on the research side that it is 100% solvable. It's just like somebody has to go do it and we haven't gotten there yet. Nobody has had the time to go and do that yet. So that's where we're at, essentially. We're working purely on large generative models for talking videos. So that's our core focus. I do think though, from a safety perspective, we have a unique framework or how we think about it. So generally videos divide into two categories. So for us, we think on one side of what is documentation, so this is the type of video that it could be a personal video where you're taking a video with your friends and you're hanging out, you're at a restaurant. It's documenting what happened. You had fun, whatever it was, it's for your memories. And there's a non-personal version of this which is like, oh, it's like a reporter documenting a crime or something that happened or whatever it is and who was involved, where was it? Maybe it was a natural disaster or something, and this is for history. We want to see what happened. And there's actually no benefit to AI-generated video in any of this. Actually, all of this, it's just negative. It's all negative. If we are generating fake versions of reality to fool people, there's just nothing good about that. And we want to stay away from that, essentially. We want to design products and build products that make it difficult to use for that particular use case, for anything that falls within that. And on the other side, you have what we think of as storytelling. Now this could be ads, it could be social media posts, it could be TV, movies. All of these things are storytelling. They're designed for entertainment, they're designed for fun. And nobody believes if you watch a Geico commercial, you're not thinking that the gecko is real selling insurance somewhere out there. You know that this is fabricated and it's for entertainment. And same with reality TV even, right? It's called reality TV. It's definitely not reality and social media, ads, all this stuff falls in the category. And if we can enable more people to tell stories and entertain other people and get their message out there, that is pure positive. This is where we want to focus. And a lot of our effort in the product and design process goes into how do we design products and build products that specifically make it really hard to use on one side and really easy to use on the other side. And that's the real challenge. That's really helpful. Something that I'm really curious about as you're chatting is ByteDance just released a really amazing model. I was actually just looking at it where you put a photo in, I think, and it just creates a video of this person talking in all these different ways. Where does that fall amongst the buckets you just described? I think that falls exactly in the area that we're in, which is talking people and that's what they're going after as well there. So that's actually one of the first examples of a large model that a larger company has released where it's able to do these dialogue or monologue videos. And I mean you yourself, you've seen it, so I'm not going to describe it too much, but as you know, it's highly expressive. It doesn't look like an avatar video. It looks like ...…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.