Joint embedding approaches, such as JEPA, started working when the focus shifted from predicting every pixel to predicting in representation space.
It started working, we abandoned this idea of predicting every pixel and basically just doing the joint embedding and predicting in representation space. That works.
Named things
JEPA · other · mentions