Meta’s Muse Voice Transcribe 70+ Languages and 20+ Speakers Simultaneously
Meta has introduced Muse Voice Transcribe, its first real-time audio model designed for live dictation and transcription throughout multiple speakers and languages.
Meta has introduced Muse Voice Transcribe, its first real-time audio model designed for live dictation and transcription throughout multiple speakers and languages.
Article outline
- What happened
- The key numbers
- What comes next
- The details
- A closer look
- The bottom line
Key points
- Muse Voice Transcribe arrives less than a week after Google introduced Gemini 3.5 Transcribe.
- Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
- Meta has priced the service at $3 per 1, 000 minutes of audio.
- According to Zuckerberg, Muse Voice Transcribe uses adaptive delay to improve transcription accuracy.
- The model was trained throughout more than 70 languages, with 25 languages validated at launch.
Meta notes the model can handle more than 20 speakers, switch between languages automatically, and understand code-switching, where residents employ words from different languages within the same sentence. Meta Pulls Scam Ads After India Uncovers Porn-Bait Banking Malware. Trained Throughout More Than 70 Languages.
Meta CEO Mark Zuckerberg demonstrated the model in a video showing it automatically identifying different speakers and switching between languages during a conversation.
According to Zuckerberg, Muse Voice Transcribe uses adaptive delay to improve transcription accuracy. It waits longer before committing to tough words while processing easier words more swiftly.
Meanwhile, the model was trained throughout more than 70 languages, with 25 languages validated at launch.
Meta notes it can additionally handle noisy and complicated real-world audio, mid-sentence language switching, and sessions lasting around an hour with more than 20 speakers.
Zuckerberg shared the demonstration after lately returning to X after roughly three years without posting on the platform.
Muse Voice Transcribe is MSL's first real-time audio perception model – rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model. Pic.twitter.com/LViMDSkbim. Mark Zuckerberg (@finkd) September 1, 2026. Meta Takes on Google Gemini.
Muse Voice Transcribe arrives less than a week after Google introduced Gemini 3.5 Transcribe. It offers similar audio transcription capabilities. Google intends to integrate its model into Android and eventually Chrome.
Meta has not remarked whether Muse Voice Transcribe will eventually become part of its major consumer services. Meta Planned to Fire Thousands Until AI Agents Failed to Deliver. Available Through Meta AI and APIs.
For now, users can access the technology through Meta's lately rolled out Meta AI Mac app.
Since the Mac app can provide voice features to other applications, Muse Voice Transcribe can additionally power dictation in other services through the app.
Developers can access the model through Muse Code and Meta's Model API.
Notably, a demonstration version of Muse Voice Transcribe is additionally available through Meta's research blog.
Muse Voice Transcribe is the latest release from Meta Superintelligence Lab (MSI). In recent weeks, the group has additionally introduced Meta's first dedicated coding agent, an open-weight model, and the Meta AI Mac app. Stay Connected with ProPakistani.
Obtain the latest tech news, telecom insights, and product launches wherever you prefer. Follow on Google Discover.
Add as a preferredSource on Google Follow on Google News Join WhatsApp.
Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and.
In short, meta's Muse Voice Transcribe 70+ Languages and 20+ Speakers Simultaneously is the central thread here, and readers can expect follow-up reporting as the picture becomes clearer.


