On Artificial Analysis’ AA-WER Streaming benchmark for English speech, Muse Voice Transcribe posted a 3.1% word error rate. Meta says that places it ahead of Cartesia Ink-2 at 3.4%, ElevenLabs Scribe v2 Real-time at 3.6%, GPT Live Transcribe at 3.9% and Gemini 3.5 Transcribe Live at 4%. Meta also reported a 17.5% diarization error rate across several public speaker-recognition benchmarks, which it says ranks first among the models included in its comparison as of Sept. 1.
The model is part of Meta’s Muse Spark family and uses an autoregressive multimodal architecture. Incoming audio is split into 80-millisecond segments, with each segment compressed into a soft token.
At every step, the model decides whether to keep listening or emit text. If it needs more context, it waits for another chunk of audio before committing to a transcription. Meta calls this approach “adaptive delay.” That design is meant to balance speed and accuracy dynamically. Easier words can be transcribed almost immediately, while more ambiguous sections receive additional audio context before the model produces text. Meta says the tradeoff is learned through reinforcement learning, combining rewards for lower word error rates with rewards for lower delay.
The same streaming architecture supports speaker diarization and endpoint detection. Muse Voice Transcribe uses special tokens to mark potential speaker changes, assign speaker identities and identify when someone begins or stops talking. Those capabilities are particularly important for conversations involving multiple people, interruptions and overlapping speech. Meta demonstrated the system with an eight-speaker conversation and says it can continue tracking speakers across much longer sessions without requiring post-processing.
Muse Voice Transcribe is also being integrated directly into Meta products. Voice dictation in Meta AI and Muse Code can now use the model, and Meta says users can invoke it across applications on Mac by holding the Fn key.
Unlike some other models in Meta’s Muse lineup, Muse Voice Transcribe will not be released with open weights.
The launch gives Meta a dedicated streaming transcription model as competition increases around low-latency speech AI. The company is emphasizing accuracy, multilingual support, speaker separation and long-context audio as the core capabilities of the new release.
This analysis is based on reporting from The New Stack & Meta.
Image courtesy of Meta.
This article was generated with AI assistance and reviewed for accuracy and quality.