Qwen-Audio-3.0-TTS is a large-scale text-to-speech (TTS) model launched by Alibaba Tongyi Qianwen, featuring a Flash version for real-time interaction and a Plus version for high-fidelity audio generation. The model has secured the top rank on the Artificial Analysis global TTS leaderboard. It supports fine-grained tag control, natural language instructions for defining voice styles, covers 16 languages and 20 Chinese dialects, delivers studio-quality audio output at up to 48KHz, and enables single synthesis of up to 3 minutes.

Fine-Grained Tag Control: Embed structured tags such as [gasp], [giggles], [angry] directly in text to precisely control intonation, emotion, and breathing nuances.
Free-Style Instruction Control: Define voice styles using natural language descriptions of character, emotion, scene, and speech rate—no professional acoustics knowledge required.
Multi-Language & Dialect Support: Covers 16 languages including Chinese, English, Japanese, Korean, German, and supports 20 Chinese dialects such as Cantonese, Chongqing, Northeastern, and Shanghai; achieves SOTA word error rates in 10 languages.
Complex Acoustic Robustness: Natively integrates voice enhancement capabilities, enabling accurate voice cloning even under high noise and high reverberation conditions.
Premium Voice Library: Offers pre-defined voice styles, dialect-specific voices, fine-grained control voices, and over 14 minority language voices—ready to use out of the box.
48KHz High-Quality Output: Upgraded from 24KHz to 48KHz, with maximum single synthesis duration of 3 minutes.
25Hz Tokenizer: Single-codebook codec emphasizing semantic content, seamlessly integrated with Qwen-Audio, enabling streaming waveform reconstruction through block-level DiT—ideal for high-quality non-streaming scenarios.
12Hz Tokenizer: 12.5Hz, 16-layer multi-codebook design enabling extreme bit-rate compression, paired with lightweight causal ConvNet for ultra-low-latency streaming reconstruction—perfect for real-time interactive applications.
Plus Version: Ideal for users seeking premium audio quality and expressive fidelity—perfect for film dubbing, audiobooks, and other high-end production scenarios.
Flash Version: Optimized for ultra-low latency real-time interaction—suitable for live streaming, chatbots, and conversational AI applications.
qwen-audio-3.0-tts-plus or qwen-audio-3.0-tts-flash service.[angry], [gasp] and Free-style natural language instructions, enabling precise control over emotion, breath, speech rate, and character portrayal.| Dimension | Qwen-Audio-3.0-TTS-Plus | ElevenLabs v3 |
|---|---|---|
| Leaderboard Ranking | Artificial Analysis #1 | Not in top 3 |
| Sampling Rate | 48KHz | Typically 44.1KHz |
| Tag Control | Native support for structured tags | Requires indirect control via Prompt |
| Dialect Support | 20 Chinese dialects | Limited |
| Multi-Language SOTA | Best WER in 10 languages | Leading in some languages only |
| Noise Robustness | Natively enhanced voice processing | Dependent on clean reference audio |
| API Pricing | $27.6 / 1M characters | Approx. $11 / 1M characters |
Disclaimer: Contains third-party opinions, does not constitute financial advice
Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi
4 days ago
StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others
4 days ago
Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series
4 days ago
Inkling – A multimodal foundational model launched by Thinking Machines Lab
5 days ago
Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360
5 days ago
X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI
6 days ago
Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model
6 days ago






