Alibaba model field guide · checked July 29, 2026

Qwen TTS turns delivery direction into speech.

Qwen-Audio-3.0-TTS is for prepared speech, not live agent dialogue. Alibaba documents natural-language style control, emotional tags, 16 languages, 20 Chinese dialects, and separate Flash and Plus tiers.

Up to 3 minutes per session 16 languages 20 Chinese dialects

Flash and Plus solve different delivery constraints

Qwen-Audio-3.0-TTS

Flash

Real-time interaction and latency-sensitive synthesis.

Time to first audio, streaming behavior, pronunciation, and accepted-output rate.

Qwen-Audio-3.0-TTS

Plus

Higher-quality generation; Alibaba documents 48 kHz studio-grade output.

Fidelity, artifacts, editing time, and whether quality improves the delivered asset.

Production review

A good sample is not a cleared asset

Test names, numbers, abbreviations, code-switching, punctuation, long pauses, emotional tags, and noisy playback. Record regeneration count, editing minutes, loudness correction, and accepted-output rate.

Confirm rights and consent for scripts, voices, likenesses, training inputs, and commercial use on the current provider and distribution surfaces. A model capability does not grant those rights.

Related routes

Official source

Alibaba Cloud's multimodal release overview documents TTS control, languages, dialects, duration, Flash, Plus, and 48 kHz output. Verify live access, pricing, limits, and rights before use.

Qwen-Audio-3.0-TTS FAQ

What is Qwen-Audio-3.0-TTS?

Qwen-Audio-3.0-TTS is Alibaba's July 2026 text-to-speech model for controllable multilingual speech, including voiceovers, audiobooks, and expressive audio.

Which languages does Qwen-Audio-3.0-TTS support?

Alibaba says it supports 16 languages and 20 Chinese dialects. Verify the current language list and test the exact accent and content you need.

What is the difference between TTS Flash and Plus?

Alibaba positions Flash for real-time interaction and Plus for higher-quality generation, including 48 kHz studio-grade output. Test both on the same scripts and delivery requirements.

How long can Qwen-Audio-3.0-TTS synthesize?

Alibaba's launch overview says up to three minutes of continuous synthesis per session. Check the live product limit before production planning.

Is Qwen-Audio-3.0-TTS available in Flowith?

This page does not claim Flowith availability. Check the live workspace model selector.