Google launched its own AI audio transcription model less than a week ago. Meta didn’t wait long to respond. The company announced Muse Voice Transcribe on September 1, its first real-time audio model, built to handle multi-speaker conversations, multiple languages, and even code-switching within a single sentence.
That last capability is worth paying attention to. Code-switching, where a speaker blends two languages mid-sentence, is notoriously difficult for transcription systems to handle cleanly. Most models either miss the switch entirely or produce garbled output. Muse Voice Transcribe is trained to detect and transcribe it accurately, which puts it ahead of most commercial alternatives on that specific use case.
Mark Zuckerberg demonstrated the model on X, showing it distinguishing between multiple speakers and switching languages in real time without manual prompting. He described the model’s core design philosophy as adaptive: it waits longer on harder words and commits faster on easier ones, using what Meta calls “adaptive delay” to balance speed and accuracy. The model is trained across 70+ languages, with 25 validated at launch, and is built to maintain accuracy across sessions with 20 or more speakers running for up to an hour.
The competitive context matters here. Google Gemini 3.5 Transcribe, announced just days earlier, targets similar capabilities. But Google has a clear integration path: Android first, Chrome eventually. Meta’s plan is less defined. Right now, Muse Voice Transcribe is accessible through the Meta AI Mac app, which can power voice features in other applications, meaning the model will handle dictation across third-party services on that platform. It’s also available to developers via Muse Code and Meta’s Model API, priced at $3 per 1,000 audio minutes.
That pricing is competitive. OpenAI’s Whisper API runs at $0.006 per minute, which works out to $6 per 1,000 minutes. AssemblyAudio and Deepgram sit in a similar range depending on the tier. Meta coming in at $3 is a meaningful undercut, especially for developers building products that process large volumes of audio. Still, pricing alone doesn’t win enterprise deals. Accuracy on real-world messy audio, latency at scale, and integration depth are what actually move the needle for production use cases.
Muse Voice Transcribe is one of several releases from Meta Superintelligence Lab in recent weeks. The group has also shipped a dedicated coding agent, an open-weight model, and the Mac app that now carries this transcription capability. The pace suggests MSI is trying to establish itself as a credible AI research and product unit, not just a branding exercise.
Key specs at launch include:
- Real-time transcription with adaptive delay for accuracy
- Support for 20+ simultaneous speakers
- 70+ languages trained, 25 validated at launch
- Mid-sentence code-switching detection
- Hour-long session support
- Priced at $3 per 1,000 audio minutes via Meta’s Model API
Whether this makes a dent depends on where Meta actually deploys it. A Mac app and a developer API are a start, but Muse Voice Transcribe needs to appear in WhatsApp, Instagram, or Ray-Ban glasses to reach the scale that would make it matter beyond the developer community. Right now, that roadmap isn’t public. And without a clear consumer integration story, this is a strong technical release that risks staying niche.




