observation · 7 Mar 2024 · 58:19

Systems based on joint embedding architectures can perform real-time speech-to-speech translation across hundreds of languages, including those without written forms.

We have systems now based on those combination of ideas that can do real time translation of hundreds of languages into each other, speech to speech. - Speech to speech, even including, which is fascinating, languages that don't have written forms- - That's right. - They're spoken only. - That's right. We don't go through text, it goes directly from speech to speech using an internal representation of kinda speech units that are discrete.

Watch at 58:19