SyncAI.news, a Varaisys broadcasting
Speaking of Voxtral
MA

Mistral AI

· 1 min read

AI LabsMistral AI

Speaking of Voxtral

Thinking

Summary

Mistral AI has launched Voxtral TTS, a 4B-parameter text-to-speech model delivering state-of-the-art multilingual voice generation with realistic, emotionally expressive speech in 9 languages, low latency, and easy voice adaptation. The model excels in contextual understanding, speaker modeling, and zero-shot cross-lingual voice adaptation, making it ideal for enterprise voice workflows and scalable AI agents. Available via API and Mistral Studio, it offers cost-effective, high-quality voice generation starting at $0.016 per 1k characters.

Today we’re releasing Voxtral TTS, our first text-to-speech model with state-of-the-art performance in multilingual voice generation. The model is lightweight at 4B parameters, making Voxtral-powered agents natural, reliable, and cost-effective at scale.

Highlights.

  1. Realistic, emotionally expressive speech in 9 popular languages with support for diverse dialects.

  2. Very low latency for time-to-first-audio.

  3. Easily adaptable to new voices.

  4. Available to test out in Mistral Studio.

  5. Enterprise-grade text-to-speech, powering critical voice agent workflows.

A natural voice generation hinges on the model’s ability to not only recite but interpret a text accurately. Contextual understanding - like neutral, happy, sarcastic, etc. - determines whether the listener considers the generation accurate or robotic. Our model excels at both contextual understanding and speaker modeling: capturing how a specific person naturally speaks. Our voice adaptation goes beyond traditional read-speech by capturing a speaker’s personality, including their natural pauses, rhythm, intonation, and emotional dexterity. With its compact size, low cost and latency, and easy adaptability, Voxtral TTS gives full control and customization for enterprises looking to own their voice AI stack.

Listen and decide: can you tell the difference?

Voice emulation

State-of-the-art performance.

Spoken natively.

Cascaded speech-to-speech translation

Prompt

Generated Audio

Voxtral TTS

Original source

This story was published by Mistral AI. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on mistral.ai

Similar News