evaluation · 7 Mar 2024 · 39:35

Yann LeCun claims that V-JEPA is the first system to learn good representations of video, enabling accurate action classification.

It's the first system that we have that learns good representations of video so that when you feed those representations to a supervised classifier head, it can tell you what action is taking place in the video with pretty good accuracy.

Watch at 39:35