Evidence receipt / belief
Published · transcript-backedJoseph Nelson: belief
7 Aug 2024 Latent Space Segment Anything 2: Demo-first Model Development
“I think so. If one wanted to be really precise with how literature refers to object re identification, object re identification is not only what SAM does for identifying that an object is similar across frames, It's also assigning a unique ID.”
Source trail
Everything needed to verify it.
- Speaker
- Joseph Nelson
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 7 Aug 2024
- Publisher
- Latent Space
Transcript context
…So I think maybe one, one thing that's different is that the object in video, it is challenging. Objects can, you know, change in appearance. There's different lighting conditions. They can deform, but I think a difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi time conversation. And so, you know, coupling this. Short term spatial memory with this, like, longer term object pointers we found was enough. So, I think that's probably one difference between vision models and LLMs. I think so. If one wanted to be really precise with how literature refers to object re identification, object re identification is not only what SAM does for identifying that an object is similar across frames, It's also assigning a unique ID. How do you think about models keeping track of occurrences of objects in addition to seeing that the same looking thing is present in multiple places? Yeah, it's a good question. I think, you know, SAM2 definitely isn't perfect and there's many limitations that, you know, we'd love to see. People in the community help us address, but one definitely challenging case is where there are multiple similar looking objects, especially if that's like a crowded scene with multiple similar looking objects, keeping track of the target object is a challenge. That's still something that I don't know if we've solved perfectly, but again, the ability to provide refinement clicks. That's one way to sort of circumvent that problem. In most cases, when there's lots of similar looking objects, if you add enough refinement clicks, you can get the perfect track throughout the video. So definitely that's one way to, to solve that problem. You know, we could have better motion estimation. We could do other things in the model to be able to disambiguate similar looking objects more effectively.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.