mms-300m: 300M-parameter speech model trained on speech data from 1,400 languages
facebook/mms-300m
About the project
This is a wav2vec2-based model pre-trained on approximately 500,000 hours of speech data across more than 1,400 languages. With 300 million parameters, it provides a foundation for speech recognition in diverse linguistic environments.
This is not a ready-to-use inference model but rather pre-trained weights intended for fine-tuning. Additional training tailored to specific tasks such as automatic speech recognition, translation, or classification is required for deployment in production services. Input speech data must strictly adhere to a 16kHz sampling rate.
The broad range of supported languages makes it suitable for building speech processing pipelines for minority languages or multilingual environments. It is licensed under CC-BY-NC 4.0, so commercial use restrictions must be checked, and it is primarily used for speech technology development for research and non-profit purposes.
facebook/mms-300m
The original page has no description.
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.