Meta released Muse Voice Transcribe on September 6, 2026, saying it transcribes speech, detects sentence boundaries, and identifies up to 20 speakers without separate systems. It is the company's first real-time audio perception model since the Spark-family model in 2025.
Meta reported a 3.1 percent word error rate on English in 0.16 seconds after a speaker finishes talking, measured on the Artificial Analysis benchmark. That compares with ElevenLabs Scribe v2 Realtime at 3.6 percent and AssemblyAI Universal-3.5 Pro Realtime at 4.0 percent.
Muse Voice Transcribe is built on reinforcement learning and targets multilingual support, handling over 70 languages and code-switching. Availability begins now, initially for Meta AI and through the Meta Model API.
"Muse Voice Transcribe breaks audio into 80-millisecond chunks and adjusts the delay for each word based on difficulty, balancing speed and accuracy on the fly," said Meta. The model dynamically controls context gathering before committing to a transcription.
The announcement follows Meta's reorganization of its AI division under Superintelligence Labs in 2025. Meta framed the release as part of its vision for personal superintelligence through AI assistants that never stop listening.
Meta did not say how many parameters the model has or its training data volume, and it isn't releasing the weights. The company said it will continue competing on pricing rather than peak performance.
Source: thedecoder