ModelSell Docs
Audio & MoreBailian Speech Synthesis

Qwen HTTP Speech Synthesis

HTTP requests and responses for Qwen3 Flash, Instruct, VC, VD and legacy Qwen-TTS.

FamilyModels
Qwen3 Flashqwen3-tts-flash, qwen3-tts-flash-2025-11-27, qwen3-tts-flash-2025-09-18
Instructqwen3-tts-instruct-flash, qwen3-tts-instruct-flash-2026-01-26
Custom voicesqwen3-tts-vc-2026-01-22, qwen3-tts-vd-2026-01-26
Legacyqwen-tts, qwen-tts-latest, qwen-tts-2025-05-22, qwen-tts-2025-04-10

Route and Parameters

POST /api/v1/services/aigc/multimodal-generation/generation
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json
FieldTypePurpose
modelstringRequired HTTP model, not a realtime model
input.textstringRequired text; legacy limit 512 tokens, Qwen3 limit 600 characters
input.voicestringRequired system or compatible custom voice
input.language_typestringOptional, e.g. Chinese, English, Auto
input.instructionsstringInstruct models only, speaking style
input.optimize_instructionsbooleanInstruct only, defaults false, requires instructions
curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/aigc/multimodal-generation/generation" \
  -H "Authorization: Bearer $MODELSELL_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3-tts-flash","input":{"text":"欢迎来到我们的语音工作室。","voice":"Cherry","language_type":"Chinese"}}'

An Instruct request:

{"model":"qwen3-tts-instruct-flash","input":{"text":"今天我们一起探索新的创意。","voice":"Cherry","language_type":"Chinese","instructions":"用平稳清晰的语气朗读,句尾稍作停顿。","optimize_instructions":false}}

For VC/VD, first clone/design a voice and pass output.voice. These models do not use Cherry; the creation target must match synthesis.

Response and Streaming

{"request_id":"REQUEST_ID","output":{"finish_reason":"stop","audio":{"url":"SIGNED_WAV_URL","data":"","id":"AUDIO_ID","expires_at":0}},"usage":{"characters":14}}

Synchronous audio is WAV. Legacy Qwen usage has input_tokens / output_tokens; Qwen3 uses characters. Add X-DashScope-SSE: enable to receive Base64 PCM in output.audio.data. Decode chunks in order; Base64 strings and raw PCM are not WAV files. Parse complete SSE events and retain final usage.

Standard speech accepts response_format: "wav" synchronously and "pcm" while streaming. The gateway does not transcode; language and optimization options go in extra_body.input.

References checked 2026-10-06: Qwen-TTS API, Model inventory.

On this page