The models target real-time voice agents, with improvements focused on streaming transcription, speech quality and response latency.
Accurate streaming transcription. Natural speech and less waiting between turns.
New MAI models
| Model | Function | Target workload |
| MAI-Transcribe-2-Streaming | Speech-to-text | Real-time transcription |
| MAI-Voice-2.1 | Text-to-speech | High-quality natural speech |
| MAI-Voice-2.1-Flash | Text-to-speech | Low-latency voice agents |
A typical implementation is:
Audio → MAI-Transcribe-2-Streaming → LLM/agent → MAI-Voice-2.1-Flash → Audio
MAI-Transcribe-2-Streaming
MAI-Transcribe-2-Streaming processes audio continuously rather than waiting for a complete recording.
This allows applications to receive partial transcripts while the user is speaking and begin downstream processing earlier.
Key features:
- streaming speech-to-text;
- incremental transcription;
- real-time processing;
- optimized use in conversational agents.
MAI-Voice-2.1 vs. MAI-Voice-2.1-Flash
Both models generate speech, but target different workloads.
| MAI-Voice-2.1 | MAI-Voice-2.1-Flash | |
| Text-to-speech | Yes | Yes |
| Natural speech | Priority | Yes |
| Low latency | Optimized | Priority |
| Voice agents | Yes | Primary target |
| Quality-focused workloads | Better fit | — |
MAI-Voice-2.1 is intended for applications prioritizing speech quality. Voice-2.1-Flash targets applications where reducing the delay before the agent starts speaking is more important.
Previous vs. new generation
| Previous | New | Main change |
| MAI transcription models | MAI-Transcribe-2-Streaming | Streaming-first transcription |
| MAI-Voice-2 | MAI-Voice-2.1 | Updated speech generation |
| MAI-Voice-2 | MAI-Voice-2.1-Flash | Dedicated low-latency variant |
The main change is the addition of models specifically optimized for streaming input and faster conversational turn-taking.
What developers should measure
For production voice agents, the important performance metrics are:
- STT first-token latency
- transcription finalization latency
- TTS time-to-first-audio
- real-time factor
- end-to-end turn latency
Microsoft has not disclosed all of these measurements publicly for the new models, so direct numerical latency comparisons with the previous generation are not yet available.
Which model to use
For a low-latency voice agent:
MAI-Transcribe-2-Streaming + LLM + MAI-Voice-2.1-Flash
For applications prioritizing speech quality:
MAI-Transcribe-2-Streaming + LLM + MAI-Voice-2.1
The new lineup separates transcription, quality-focused speech generation and latency-focused speech generation into dedicated models, allowing developers to select the voice stack based on application requirements.