ModelSell Docs
Audio & MoreBailian Speech Synthesis

Qwen-Audio and CosyVoice Synthesis

SpeechSynthesizer HTTP, SSE and inference WebSocket parameters and outputs.

Models: qwen-audio-3.0-tts-plus, qwen-audio-3.1-tts-flash, qwen-audio-3.0-tts-flash, cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-plus, cosyvoice-v3-flash, cosyvoice-v2, cosyvoice-v1.

HTTP Parameters

POST /api/v1/services/audio/tts/SpeechSynthesizer
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json

CosyVoice v1 requires inference WS or the standard speech bridge.

FieldTypePurpose
modelstringRequired model ID
input.text / input.voicestringRequired text and compatible voice
input.format / input.sample_ratestring / integerCodec and sample rate
input.volume / rate / pitchnumberVolume, speed, pitch
input.instruction / language_hintsstring / string[]Supported style or language controls
input.enable_ssml / word_timestamp_enabledbooleanSSML and streaming word timestamps
input.seed / hot_fixinteger / objectSeed and pronunciation/replacement rules

Options and ranges depend on the model and voice. Tested voices include longanhuan_v3.1 for Qwen-Audio 3.1, longanhuan_v3.6 for Qwen-Audio 3.0 Flash, longanyang for CosyVoice v3, and longxiaochun for v1. For v2/v3.5, create a compatible custom voice.

curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/audio/tts/SpeechSynthesizer" \
  -H "Authorization: Bearer $MODELSELL_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "cosyvoice-v3-flash",
    "input": {"text": "欢迎收听本期节目。", "voice": "longanyang", "format": "mp3", "sample_rate": 24000, "rate": 1.0}
  }' --output response.json

Response and SSE

{"request_id":"REQUEST_ID","output":{"finish_reason":"stop","audio":{"url":"SIGNED_AUDIO_URL","data":"","id":"AUDIO_ID","expires_at":0}},"usage":{"characters":9}}

Download output.audio.url for complete audio. Qwen-Audio 3.1 uses usage.input_tokens / output_tokens; other listed models use characters.

Add X-DashScope-SSE: enable and use curl -N. Parse events at blank-line boundaries; join all data: lines of one event. Decode output.audio.data as Base64 in order. Sentence events and output.sentence.words preserve timestamps. Do not sum repeated cumulative usage.

Duplex WebSocket

Connect to $MODELSELL_WS_URL/api-ws/v1/inference?model=cosyvoice-v3-flash with Bearer authentication. The query and payload.model must match.

  1. Send run-task with a fresh UUID.
  2. Wait for task-started before sending continue-task text.
  3. Send finish-task; read audio until task-finished.
  4. Handle task-failed. Reuse the connection with a new task ID and the same model.
{
  "header": {"action": "run-task", "task_id": "NEW_UUID", "streaming": "duplex"},
  "payload": {
    "task_group": "audio", "task": "tts", "function": "SpeechSynthesizer",
    "model": "cosyvoice-v3-flash", "input": {},
    "parameters": {"text_type": "PlainText", "voice": "longanyang", "format": "mp3", "sample_rate": 24000}
  }
}

WS synthesis options belong in payload.parameters, not HTTP input:

{"header":{"action":"continue-task","task_id":"NEW_UUID","streaming":"duplex"},"payload":{"input":{"text":"欢迎收听本期节目。"}}}
{"header":{"action":"finish-task","task_id":"NEW_UUID","streaming":"duplex"},"payload":{"input":{}}}

Audio arrives as binary frames, not JSON/Base64. With SSML enabled, send only one continue-task. Newer options are not universally supported by v1.

References checked 2026-10-06: Qwen-Audio HTTP, CosyVoice HTTP, WS client events, Realtime synthesis.

On this page