Meta recently developed Muse Voice Transcribe, which is a real-time audio perception model that is available for five major Indian languages, such as Hindi, Tamil, Telugu, Kannada, and Malayalam. This model was built in Meta Superintelligence Labs, and it is the first real-time audio perception model developed by Meta, which supports more than 70 languages in total.
Meta Muse Voice Transcribe brings real-time speech transcription, speaker separation, and endpointing to one solution. The model is able to work in real-time mode with the ability to differentiate between different speakers and detect start/end speaking moments.
Meta Muse Voice Transcribe Offers Real-Time Speech Transcription
As stated by Meta, Muse Voice Transcribe is capable of transcribing speech to text in real time, when the person is speaking, rather than waiting for the entire audio recording to be completed before transcribing it.
The AI model has the ability to identify more than 20 speakers in a single recording and can handle recordings that go beyond an hour in length.
Meta Muse Voice Transcribe Processes Audio in 80-Millisecond Chunks
Muse Voice Transcribe, according to Meta, operates on 80-millisecond audio segments at a time and determines when it has collected sufficient data to accurately transcribe individual words.
This algorithm is given more time to transcribe more difficult words and less time to process easier words. It has been noted that reinforcement learning plays an important role in this process.
Meta Muse Voice Transcribe Supports Five Indian Languages
Meta trained Muse Voice Transcribe on over 70 languages and tested 25 languages thoroughly in preparation for its first release.
Languages supported by the model in India include the following:
| Language Support | Details |
|---|---|
| Hindi | Supported for real-time transcription |
| Tamil | Supported for real-time transcription |
| Telugu | Supported for real-time transcription |
| Kannada | Supported for real-time transcription |
| Malayalam | Supported for real-time transcription |
Additionally, the system can be used in multilingual scenarios, where the speaker may shift from one language to another even within a sentence itself.
Moreover, Muse Voice Transcribe uses language, keyword, and context biasing techniques, which assist in recognizing words based on the audio and the whole conversation.
Meta Muse Voice Transcribe Availability and Pricing
Meta has released Muse Voice Transcribe via the Meta Model API for developers to utilise the model for transcribing speech.
This model is currently being applied in the dictation of Meta AI for Mac and Muse Code.
The cost of the API is US$3 (INR 300) for every 1,000 audio minutes, roughly equivalent to US$0.18 (INR 17) per hour.
Meta Muse Voice Transcribe adds multilingual support for Indian languages, making it more useful for developers to build voice-based applications.
Stay connected with the latest technology updates, digital trends, and industry stories on SmartMag.


