Qwen HTTP Speech Synthesis
HTTP requests and responses for Qwen3 Flash, Instruct, VC, VD and legacy Qwen-TTS.
| Family | Models |
|---|---|
| Qwen3 Flash | qwen3-tts-flash, qwen3-tts-flash-2025-11-27, qwen3-tts-flash-2025-09-18 |
| Instruct | qwen3-tts-instruct-flash, qwen3-tts-instruct-flash-2026-01-26 |
| Custom voices | qwen3-tts-vc-2026-01-22, qwen3-tts-vd-2026-01-26 |
| Legacy | qwen-tts, qwen-tts-latest, qwen-tts-2025-05-22, qwen-tts-2025-04-10 |
Route and Parameters
POST /api/v1/services/aigc/multimodal-generation/generation
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json| Field | Type | Purpose |
|---|---|---|
model | string | Required HTTP model, not a realtime model |
input.text | string | Required text; legacy limit 512 tokens, Qwen3 limit 600 characters |
input.voice | string | Required system or compatible custom voice |
input.language_type | string | Optional, e.g. Chinese, English, Auto |
input.instructions | string | Instruct models only, speaking style |
input.optimize_instructions | boolean | Instruct only, defaults false, requires instructions |
curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/aigc/multimodal-generation/generation" \
-H "Authorization: Bearer $MODELSELL_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3-tts-flash","input":{"text":"欢迎来到我们的语音工作室。","voice":"Cherry","language_type":"Chinese"}}'An Instruct request:
{"model":"qwen3-tts-instruct-flash","input":{"text":"今天我们一起探索新的创意。","voice":"Cherry","language_type":"Chinese","instructions":"用平稳清晰的语气朗读,句尾稍作停顿。","optimize_instructions":false}}For VC/VD, first clone/design a voice and pass output.voice. These models do not use Cherry; the creation target must match synthesis.
Response and Streaming
{"request_id":"REQUEST_ID","output":{"finish_reason":"stop","audio":{"url":"SIGNED_WAV_URL","data":"","id":"AUDIO_ID","expires_at":0}},"usage":{"characters":14}}Synchronous audio is WAV. Legacy Qwen usage has input_tokens / output_tokens; Qwen3 uses characters. Add X-DashScope-SSE: enable to receive Base64 PCM in output.audio.data. Decode chunks in order; Base64 strings and raw PCM are not WAV files. Parse complete SSE events and retain final usage.
Standard speech accepts response_format: "wav" synchronously and "pcm" while streaming. The gateway does not transcode; language and optimization options go in extra_body.input.
References checked 2026-10-06: Qwen-TTS API, Model inventory.