Today, speech technology is only available for a small fraction of the thousands of languages spoken around the world because traditional systems need to be trained on large amounts of annotated speech audio with transcriptions. Obtaining that kind of data for every human language and dialect is almost impossible.
Wav2vec works around this limitation by requiring little to no transcribed data. The model uses self-supervision to push the boundaries by learning from unlabeled training data. This enables speech recognition systems for many more languages and dialects, such as Kyrgyz and Swahili, which don’t have a lot of transcribed speech audio. Self-supervision is the key to leveraging unannotated data and building better systems.
P. Moreira, Y. Bizzoni, K. Nielbo, I. Lassen, и M. Thomsen. Proceedings of the The 5th Workshop on Narrative Understanding, стр. 25--35. Toronto, Canada, Association for Computational Linguistics, (июля 2023)
A. Blom, F. Carlsson, и E. Wihlborg. Proceedings of the 55th Hawaii International Conference on System Sciences | 2022, стр. 2563-2572. Honolulu, (2022)(Eurobarometer).
T. Piske, и A. Steinlen. Cognition and Second Language Acquisition: Studies on pre-school, primary school and secondary school children, том 4 из Multilingualism and Language Teaching, Narr Francke Attempto Verlag, Tübingen, (Mikrozensus).(2022)
O. Decker, A. Yendell, A. Heller, и E. Brähler. Autoritäre Dynamiken in unsicheren Zeiten. Neue Herausforderungen - alte Reaktionen? / Leipziger Autoritarismus Studie 2022, Psychosozial-Verlag, Gießen, (ALLBUS).(2022)