Xiaomi’s New Open-Source AI Can Isolate One Speaker From Overlapping Voices

Xiaomi has published and open-sourced Xiaomi-CocktailASR-1, a speech recognition model designed to identify and transcribe one person's voice even when a number of individuals are speaking meanwhile.

TechnologyNews Info Wire3 min read
Xiaomi’s New Open-Source AI Can Isolate One Speaker From Overlapping Voices

Xiaomi has published and open-sourced Xiaomi-CocktailASR-1, a speech recognition model designed to identify and transcribe one person's voice even when a number of individuals are speaking meanwhile.

Article outline

  1. What happened
  2. The details
  3. Why it matters
  4. A closer look
  5. More on the story
  6. The bottom line

Key points

  • Xiaomi Pad 9 and Pad 9 Pro Arrive Shortly After Pad 9 Pro Max.
  • Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
  • Developers can now access Xiaomi-CocktailASR-1 through GitHub and Hugging Face for testing and further development.
  • Add as a preferredSource on Google Follow on Google News Join WhatsApp.
  • CocktailASR-1 works by first taking a short audio sample of the person a user wants to follow.

Xiaomi developed the model to address what is commonly known as the "cocktail party problem, " where overlapping voices can confuse conventional automatic speech recognition systems.

Xiaomi Pad 9 and Pad 9 Pro Arrive Shortly After Pad 9 Pro Max. It Can Recognize Voices.

Meanwhile, the model then uses that sample as a voice reference while processing a recording containing multiple speakers. It attempts to identify only the selected person's speech and transcribe it while ignoring the others.

This could be useful for meetings, interviews, group conversations, and other recordings where a number of residents speak over one another. Humanoid Robots May Finally Be Ready for Homes. Built Around an LLM Architecture. Xiaomi notes CocktailASR-1 uses an end-to-end substantial language model architecture.

According to the firm, the model achieved state-of-the-art results throughout a number of multi-speaker speech recognition benchmarks and outperformed existing systems designed for similar tasks.

Xiaomi additionally states the model maintains competitive performance in recordings containing only one speaker, meaning the multi-speaker capabilities do not significantly reduce standard transcription quality. Avoids Transcribing the Wrong Person. The model is additionally designed to avoid transcribing the wrong person.

If the selected speaker is not present in a recording, CocktailASR-1 can return an empty result instead of attempting to assign another person's speech to the target speaker.

Notably, the model additionally includes a reasoning mode that can provide further information associated with how it arrived at a transcription. Available on GitHub and Hugging Face.

CocktailASR-1 is the latest AI model Xiaomi has created available to developers. It follows other open-source releases, including MiMo-V2-Flash and Xiaomi Robotics-0.

Developers can now access Xiaomi-CocktailASR-1 through GitHub and Hugging Face for testing and further development. Stay Connected with ProPakistani.

Obtain the latest tech news, telecom insights, and product launches wherever you prefer. Follow on Google Discover.

Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and.

For now, xiaomi's New Open-Source AI Can Isolate One Speaker From Overlapping Voices remains the part of the story worth watching, and further updates are likely as more details are confirmed.

Leave a Reply

Your email address will not be published. Required fields are marked *