Qwen-Audio-3.0-TTS – A Text-to-Speech (TTS) model launched by Alibaba Tongyi Qianwen

Qwen-Audio-3.0-TTS – A Text-to-Speech (TTS) model launched by Alibaba Tongyi Qianwen

What is Qwen-Audio-3.0-TTS

Qwen-Audio-3.0-TTS is a large-scale text-to-speech (TTS) model launched by Alibaba Tongyi Qianwen, featuring a Flash version for real-time interaction and a Plus version for high-fidelity audio generation. The model has secured the top rank on the Artificial Analysis global TTS leaderboard. It supports fine-grained tag control, natural language instructions for defining voice styles, covers 16 languages and 20 Chinese dialects, delivers studio-quality audio output at up to 48KHz, and enables single synthesis of up to 3 minutes.

Qwen-Audio-3.0-TTS

Main Features of Qwen-Audio-3.0-TTS

  • Fine-Grained Tag Control: Embed structured tags such as [gasp], [giggles], [angry] directly in text to precisely control intonation, emotion, and breathing nuances.

  • Free-Style Instruction Control: Define voice styles using natural language descriptions of character, emotion, scene, and speech rate—no professional acoustics knowledge required.

  • Multi-Language & Dialect Support: Covers 16 languages including Chinese, English, Japanese, Korean, German, and supports 20 Chinese dialects such as Cantonese, Chongqing, Northeastern, and Shanghai; achieves SOTA word error rates in 10 languages.

  • Complex Acoustic Robustness: Natively integrates voice enhancement capabilities, enabling accurate voice cloning even under high noise and high reverberation conditions.

  • Premium Voice Library: Offers pre-defined voice styles, dialect-specific voices, fine-grained control voices, and over 14 minority language voices—ready to use out of the box.

  • 48KHz High-Quality Output: Upgraded from 24KHz to 48KHz, with maximum single synthesis duration of 3 minutes.

Technical Principles of Qwen-Audio-3.0-TTS

  • Dual-Track Hybrid Streaming Architecture: Employs a Dual-Track LM architecture enabling both streaming and non-streaming generation within a single model. Inputting a single character immediately triggers the first audio packet, achieving end-to-end synthesis latency as low as 97ms.
  • Discrete Multi-Codebook Language Model: Implements full-information end-to-end speech modeling via a discrete multi-codebook LM, bypassing information bottlenecks and cascaded errors inherent in traditional LM+DiT pipelines, significantly enhancing generalization and generation efficiency.
  • Dual Tokenizer Design:
    • 25Hz Tokenizer: Single-codebook codec emphasizing semantic content, seamlessly integrated with Qwen-Audio, enabling streaming waveform reconstruction through block-level DiT—ideal for high-quality non-streaming scenarios.

    • 12Hz Tokenizer: 12.5Hz, 16-layer multi-codebook design enabling extreme bit-rate compression, paired with lightweight causal ConvNet for ultra-low-latency streaming reconstruction—perfect for real-time interactive applications.

  • Large-Scale Multilingual Training: Trained on over 5 million hours of diverse multilingual speech data across 10 languages, supporting 3-second rapid voice cloning and text-based voice design (Voice Design).
  • Instruction-Driven Acoustic Control: Treats voice control as a language modeling task. Uses natural language instructions in ChatML format to flexibly manipulate multidimensional acoustic attributes such as voice timbre, emotion, speech rate, and character identity—achieving “what you hear is what you think”.

How to Use Qwen-Audio-3.0-TTS

  • Select Version:
    • Plus Version: Ideal for users seeking premium audio quality and expressive fidelity—perfect for film dubbing, audiobooks, and other high-end production scenarios.

    • Flash Version: Optimized for ultra-low latency real-time interaction—suitable for live streaming, chatbots, and conversational AI applications.

  • Integrate via Platform: Access the Alibaba Cloud BaiLian console, search, and enable either the qwen-audio-3.0-tts-plus or qwen-audio-3.0-tts-flash service.
  • Call API: Submit text via standard API, optionally including structured tags or natural language instructions, to receive 48KHz synthesized audio output.
  • Choose Voice: Select from the premium voice library’s preset voices, or upload reference audio for voice cloning.

Core Advantages of Qwen-Audio-3.0-TTS

  • Leaderboard Performance: Ranked #1 globally on the Artificial Analysis TTS leaderboard; achieved perfect scores across 16 languages in speaker similarity evaluation (Plus version), and SOTA word error rates in 10 languages.
  • Comprehensive Dual-Version Coverage: Flash version offers 97ms first-packet latency for real-time interactivity; Plus version delivers 48KHz output for cinematic-grade audio quality.
  • Fine-Grained Expressiveness: Supports structured tags like [angry], [gasp] and Free-style natural language instructions, enabling precise control over emotion, breath, speech rate, and character portrayal.
  • Multi-Language & Dialect Support: Covers 16 languages and 20 Chinese dialects; specialized training mitigates “dialect accent dilution,” faithfully reproducing native-speaker authenticity.
  • Acoustic Robustness: Natively embedded voice enhancement ensures accurate voice cloning even in noisy or highly reverberant environments—no need for quiet recording conditions.
  • Ultra-Fast Cloning: Complete voice cloning in just 3 seconds with reference audio, supports text-driven voice design (Voice Design).

Project Repository

  • Official Website: https://funaudiollm.github.io/qwen-audio-3.0-tts/

Competitive Comparison with Similar Models

Dimension Qwen-Audio-3.0-TTS-Plus ElevenLabs v3
Leaderboard Ranking Artificial Analysis #1 Not in top 3
Sampling Rate 48KHz Typically 44.1KHz
Tag Control Native support for structured tags Requires indirect control via Prompt
Dialect Support 20 Chinese dialects Limited
Multi-Language SOTA Best WER in 10 languages Leading in some languages only
Noise Robustness Natively enhanced voice processing Dependent on clean reference audio
API Pricing $27.6 / 1M characters Approx. $11 / 1M characters

Application Scenarios

  • Film, TV & Game Dubbing: Precisely control emotional arcs and breathing rhythms via structured tags, enabling character-driven, performance-level TTS ideal for animation, television dramas, and video game NPCs.
  • Audiobooks & Podcasts: Define narration style and pacing using natural language instructions. Deliver high-quality 48KHz audio output up to 3 minutes per batch—supports mass production of long-form audio content.
  • Intelligent Customer Service & Assistants: Flash version’s 97ms ultra-low first-packet latency combined with support for 20 Chinese dialects enables natural, fluent real-time voice interaction—significantly improving user experience.
  • Online Education: Clone instructor voices and adjust speech rate and emotion to generate personalized teaching audio, supporting localized voice production for multilingual course content.
  • Live Streaming & Real-Time Interaction: Flash version’s low-latency capability suits real-time caption reading, virtual streamer voicing, and live e-commerce scenarios requiring instant audio feedback.
#AI Tools

Disclaimer: Contains third-party opinions, does not constitute financial advice

Share To
X
Telegram
WeChat
QQ
Link
Recommended Reading

Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi

4 days ago
Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi

StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others

4 days ago
StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others

Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series

4 days ago
Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series

Inkling – A multimodal foundational model launched by Thinking Machines Lab

5 days ago
Inkling – A multimodal foundational model launched by Thinking Machines Lab

Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360

5 days ago
Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360

X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI

6 days ago
X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI

Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model

6 days ago
Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model