Evidence receipt / commitment
Published · transcript-backedGaurav Misra: commitment
27 Mar 2025 Lenny's Podcast How to win in the AI era: Ship a feature every week, embrace technical debt, ruthlessly cut scope, and create magic your competitors can't copy | Gaurav Misra (CEO and co-founder of Captions)
“So generally videos divide into two categories. So for us, we think on one side of what is documentation, so this is the type of video that it could be a personal video where you're taking a video with your friends and you're hanging out, you're at a restaurant.”
Source trail
Everything needed to verify it.
- Speaker
- Gaurav Misra
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 27 Mar 2025
- Publisher
- Lenny's Podcast
Transcript context
…It's fundamental. At the end of the day, a time where video images, audio can't be trusted actually hasn't existed for a while. If you think about ... I mean there was a world in the 1800s where there was no video or audio or images and everything was proven by he said, she said for the most part. And it's possible that if everything can be generated and anything can be created and it looks just as real as if it were real and there's no way to tell, then we might actually return to that world where there's no way to prove anything besides physical evidence or he said, she said. And I think that's scary, but also possibly opens a bunch of new opportunity for someone to figure out how to solve this problem. I think it's going to be a big problem. I do think today, we are almost there in terms of creating absolutely photorealistic video. I mean the very recent models, a very cutting edge is just about ... It feels like a few centimeters away from achieving it, but I do think to fully get there to the point where it cannot be differentiated at all, it's still a couple of years away. I also think that it is use case driven in a way. I think thinking about Captions for a second, we take a unique view on what type of video we want to focus on. Video generation and text to video generation. If you look at it today, it's all silent video. There's no audio and it's often what you think of as stop video or B-roll, right? You can actually make a movie with B-roll. And a lot of a movie or a TV show or a social media post or an ad actually is dialogue or monologue. That's actually what it is is people talking to each other, to the camera, interacting. That's actually what makes true story. B-roll is supportive elements that are showing up to set the scene or something like maybe before the scene opens, you see a few shots of New York City or LA or something, and then you jump into the room and now two people are talking. So our goal is to solve the talking video problem. How do we create video where people are delivering dialogue or monologue or things like that? And that's what we focus on purely. And there actually isn't a lot of work happening in that area today and it's not a solved problem. We're getting there, we're getting closer and closer, but today's models actually bifurcate a little bit. So there's a set of companies today that are able to create these types of what we're talking about is avatar videos. They're using this technology called neural rendering. It's actually not a technology that's affected by the transformer and diffusion model revolution or the large model revolution, essentially. This is a technology that existed separately and it doesn't have anything to do with the AI growth happening right now. It just happens to produce semi-realistic outputs, but it actually stops at some point because it's not clear how it becomes generalizable in every situation. It has to be trained on people individually. just happens to produce semi-realistic outputs, but it actually stops at some point because it's not clear how it becomes generalizable in every situation. It has to be trained on people individually. So you might ingest a little bit of video of you and then you can generate you. And so it's a different technology and a different outcome, essentially. And a bunch of companies using this type of model, a bunch of companies are doing general text to video with no audio today. These are large generative models and they have the capability to do more, but that frontier just hasn't been reached yet. I think there's no doubt in anybody's mind on the research side that it is 100% solvable. It's just like somebody has to go do it and we haven't gotten there yet. Nobody has had the time to go and do that yet. So that's where we're at, essentially. We're working purely on large generative models for talking videos. So that's our core focus. I do think though, from a safety perspective, we have a unique framework or how we think about it. So generally videos divide into two categories. So for us, we think on one side of what is documentation, so this is the type of video that it could be a personal video where you're taking a video with your friends and you're hanging out, you're at a restaurant. It's documenting what happened. You had fun, whatever it was, it's for your memories. And there's a non-personal version of this which is like, oh, it's like a reporter documenting a crime or something that happened or whatever it is and who was involved, where was it? Maybe it was a natural disaster or something, and this is for history. We want to see what happened. And there's actually no benefit to AI-generated video in any of this. Actually, all of this, it's just negative. It's all negative. If we are generating fake versions of reality to fool people, there's just nothing good about that. And we want to stay away from that, essentially. We want to design products and build products that make it difficult to use for that particular use case, for anything that falls within that. And on the other side, you have what we think of as storytelling. Now this could be ads, it could be social media posts, it could be TV, movies. All of these things are storytelling. They're designed for entertainment, they're designed for fun. And nobody believes if you watch a Geico commercial, you're not thinking that the gecko is real selling insurance somewhere out there. You know that this is fabricated and it's for entertainment. And same with reality TV even, right? It's called reality TV. It's definitely not reality and social media, ads, all this stuff falls in the category. And if we can enable more people to tell stories and entertain other people and get their message out there, that is pure positive. This is where we want to focus. And a lot of our effort in the product and design process goes into how do we design products and build products that specifically make it really hard to use on one side and really easy to use on the other side. And that's the real challenge. That's really helpful. Something that I'm really curious about as you're chatting is ByteDance just released a really amazing model. I was actually just looking at it where you put a photo in, I think, and it just creates a video of this person talking in all these different ways. Where does that fall amongst the buckets you just described?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.