Qwen-Audio and CosyVoice Synthesis
SpeechSynthesizer HTTP, SSE and inference WebSocket parameters and outputs.
Models: qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash, cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-plus, cosyvoice-v3-flash, cosyvoice-v2, cosyvoice-v1.
HTTP Parameters
POST /api/v1/services/audio/tts/SpeechSynthesizer
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/jsonCosyVoice v1 requires inference WS or the standard speech bridge.
| Field | Type | Purpose |
|---|---|---|
model | string | Required model ID |
input.text / input.voice | string | Required text and compatible voice |
input.format / input.sample_rate | string / integer | Codec and sample rate |
input.volume / rate / pitch | number | Volume, speed, pitch |
input.instruction / language_hints | string / string[] | Supported style or language controls |
input.enable_ssml / word_timestamp_enabled | boolean | SSML and streaming word timestamps |
input.seed / hot_fix | integer / object | Seed and pronunciation/replacement rules |
Options and ranges depend on the model and voice. Tested voices include longanhuan_v3.1 for Qwen-Audio 3.1, longanhuan_v3.6 for Qwen-Audio 3.0 Flash, longanyang for CosyVoice v3, and longxiaochun for v1. For v2/v3.5, create a compatible custom voice.
curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/audio/tts/SpeechSynthesizer" \
-H "Authorization: Bearer $MODELSELL_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "cosyvoice-v3-flash",
"input": {"text": "欢迎收听本期节目。", "voice": "longanyang", "format": "mp3", "sample_rate": 24000, "rate": 1.0}
}' --output response.jsonResponse and SSE
{"request_id":"REQUEST_ID","output":{"finish_reason":"stop","audio":{"url":"SIGNED_AUDIO_URL","data":"","id":"AUDIO_ID","expires_at":0}},"usage":{"characters":9}}Download output.audio.url for complete audio. Qwen-Audio 3.1 uses usage.input_tokens / output_tokens; other listed models use characters.
Add X-DashScope-SSE: enable and use curl -N. Parse events at blank-line boundaries; join all data: lines of one event. Decode output.audio.data as Base64 in order. Sentence events and output.sentence.words preserve timestamps. Do not sum repeated cumulative usage.
Duplex WebSocket
Connect to $MODELSELL_WS_URL/api-ws/v1/inference?model=cosyvoice-v3-flash with Bearer authentication. The query and payload.model must match.
- Send
run-taskwith a fresh UUID. - Wait for
task-startedbefore sendingcontinue-tasktext. - Send
finish-task; read audio untiltask-finished. - Handle
task-failed. Reuse the connection with a new task ID and the same model.
{
"header": {"action": "run-task", "task_id": "NEW_UUID", "streaming": "duplex"},
"payload": {
"task_group": "audio", "task": "tts", "function": "SpeechSynthesizer",
"model": "cosyvoice-v3-flash", "input": {},
"parameters": {"text_type": "PlainText", "voice": "longanyang", "format": "mp3", "sample_rate": 24000}
}
}WS synthesis options belong in payload.parameters, not HTTP input:
{"header":{"action":"continue-task","task_id":"NEW_UUID","streaming":"duplex"},"payload":{"input":{"text":"欢迎收听本期节目。"}}}{"header":{"action":"finish-task","task_id":"NEW_UUID","streaming":"duplex"},"payload":{"input":{}}}Audio arrives as binary frames, not JSON/Base64. With SSML enabled, send only one continue-task. Newer options are not universally supported by v1.
References checked 2026-10-06: Qwen-Audio HTTP, CosyVoice HTTP, WS client events, Realtime synthesis.