observation · 7 Mar 2024 · 33:17

Perceptual inputs like vision contain far more redundancy than text, making self-supervised learning more effective for such modalities.

there is way more redundancy in the structure in perceptual inputs, sensory input like vision, than there is in text, which is not nearly as redundant.

Watch at 33:17