Yann LeCun describes V-JEPA as an extension of I-JEPA applied to video, where a temporal tube is masked across frames.
A more recent version of this that we have is called V-JEPA. So it's basically the same idea as I-JEPA except it's applied to video. So now you take a whole video and you mask a whole chunk of it.