Evidence receipt / belief
Published · transcript-backedNikhila Ravi: belief
7 Aug 2024 Latent Space Segment Anything 2: Demo-first Model Development
“There's different lighting conditions. They can deform, but I think a difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi time conversation.”
Source trail
Everything needed to verify it.
- Speaker
- Nikhila Ravi
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 7 Aug 2024
- Publisher
- Latent Space
Transcript context
…Amazing. My main reaction to finding out about the context length of eight input frames and six pass frames as their default is why not 60? Why not 600? In text language models, we're very used to severely extending context windows. And what does that do to the memory of your model? So I think maybe one, one thing that's different is that the object in video, it is challenging. Objects can, you know, change in appearance. There's different lighting conditions. They can deform, but I think a difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi time conversation. And so, you know, coupling this. Short term spatial memory with this, like, longer term object pointers we found was enough. So, I think that's probably one difference between vision models and LLMs. I think so. If one wanted to be really precise with how literature refers to object re identification, object re identification is not only what SAM does for identifying that an object is similar across frames, It's also assigning a unique ID. How do you think about models keeping track of occurrences of objects in addition to seeing that the same looking thing is present in multiple places?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.