Misc,

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Y. Zhang, W. Han, J. Qin, Y. Wang, A. Bapna, Z. Chen, N. Chen, B. Li, V. Axelrod, G. Wang, Z. Meng, K. Hu, A. Rosenberg, R. Prabhavalkar, D. Park, P. Haghani, J. Riesa, G. Perng, H. Soltau, T. Strohman, B. Ramabhadran, T. Sainath, P. Moreno, C. Chiu, J. Schalkwyk, F. Beaufays, and Y. Wu.
(2023)cite arxiv:2303.01037Comment: 20 pages, 7 figures, 8 tables.

Abstract

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the encoder of the model on a large unlabeled multilingual dataset of 12 million (M) hours spanning over 300 languages, and fine-tuning on a smaller labeled dataset. We use multilingual pre-training with random-projection quantization and speech-text modality matching to achieve state-of-the-art performance on downstream multilingual ASR and speech-to-text translation tasks. We also demonstrate that despite using a labeled training set 1/7-th the size of that used for the Whisper model, our model exhibits comparable or better performance on both in-domain and out-of-domain speech recognition tasks across many languages.

BibTeX key: zhang2023google
entry type: misc
year: 2023
url: http://arxiv.org/abs/2303.01037
note: cite arxiv:2303.01037Comment: 20 pages, 7 figures, 8 tables

BibSonomy

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Abstract

Tags

Users

Comments and Reviewsshow / hide

Cite this publication

More citation styles

search on