observation · 7 Mar 2024 · 57:55

Wav2Vec, a joint embedding architecture trained with contrastive learning, enables multilingual speech recognition systems with mostly unlabeled data and minimal labeled data.

We had similar success in speech recognition, a system called Wav2Vec, which is also a joint embedding architecture by the way, trained with contrastive learning. And that system also can produce speech recognition systems that are multilingual with mostly unlabeled data and only need a few minutes of labeled data to actually do speech recognition.

Watch at 57:55

Named things

Wav2Vec · tool · mentions