Voxtral Mini Transcribe 2
A batch speech-to-text API from Mistral AI that helps developers and teams turn multilingual audio into transcripts with speaker labels, timestamps, and context biasing.
Tool overview
Based on the available evidence, Voxtral Mini Transcribe 2 looks promising, but it is better described as a newly well-publicized API than a broadly validated default choice. The adoption judgment is cautiously positive: official launch posts repeatedly highlight batch transcription price-performance, 4% WER on FLEURS, $0.003/min, plus diarization, word-level timestamps, and context biasing. Those are useful signals, but they are still launch claims and attention proof, not the same as broad production proof.
Its practical role is also easy to misunderstand. It is not a general voice assistant, not an open-weight self-hosted ASR stack, and not the realtime model in the Voxtral family. A better comparison is a production-oriented batch transcription API, closer to Whisper API or Deepgram batch transcription. For usability proof, the evidence is narrower but more valuable: one Zhihu write-up tested a roughly 45-minute noisy multilingual meeting and reported good speaker separation and solid terminology recognition; on X, developers compared it against local whisper-large-v3-turbo and WhisperX on Replicate, which suggests it is competitive enough to benchmark seriously.