ModelSell Docs
Audio & MoreBailian Speech Synthesis

Bailian TTS Models and Integration

API, format, voice and billing map for 36 synthesis models and 3 voice services.

These guides cover Qwen-Audio-TTS, CosyVoice, Qwen-TTS and MiniMax. The documented inventory was checked on 2026-10-06. Model availability, token permissions and current prices depend on your ModelSell group.

Authentication

Use a ModelSell API key and the ModelSell base URL. The channel manages provider credentials and workspace addresses.

export MODELSELL_BASE_URL="https://api.modelsell.com"
export MODELSELL_WS_URL="wss://api.modelsell.com"
export MODELSELL_API_KEY="YOUR_MODELSELL_API_KEY"

HTTP uses Authorization: Bearer $MODELSELL_API_KEY and Content-Type: application/json. WebSocket handshakes use the same Authorization header.

Complete Model and Documentation Map

Model IDIntegration guideProtocolUsage unit
qwen-audio-3.0-tts-plusQwen-Audio / CosyVoiceHTTP / inference WSCharacters
qwen-audio-3.1-tts-flashQwen-Audio / CosyVoiceHTTP / inference WSInput/output tokens
qwen-audio-3.0-tts-flashQwen-Audio / CosyVoiceHTTP / inference WSCharacters
cosyvoice-v3.5-plusQwen-Audio / CosyVoiceHTTP / inference WSCharacters
cosyvoice-v3.5-flashQwen-Audio / CosyVoiceHTTP / inference WSCharacters
cosyvoice-v3-plusQwen-Audio / CosyVoiceHTTP / inference WSCharacters
cosyvoice-v3-flashQwen-Audio / CosyVoiceHTTP / inference WSCharacters
cosyvoice-v2Qwen-Audio / CosyVoiceHTTP / inference WSCharacters
cosyvoice-v1Qwen-Audio / CosyVoiceinference WS;speech bridgeCharacters
qwen3-tts-flashQwen HTTPHTTPCharacters
qwen3-tts-flash-2025-11-27Qwen HTTPHTTPCharacters
qwen3-tts-flash-2025-09-18Qwen HTTPHTTPCharacters
qwen3-tts-instruct-flashQwen HTTPHTTPCharacters
qwen3-tts-instruct-flash-2026-01-26Qwen HTTPHTTPCharacters
qwen3-tts-vc-2026-01-22Qwen HTTP · CloningHTTPCharacters
qwen3-tts-vd-2026-01-26Qwen HTTP · DesignHTTPCharacters
qwen-ttsQwen HTTPHTTPInput/output tokens
qwen-tts-latestQwen HTTPHTTPInput/output tokens
qwen-tts-2025-05-22Qwen HTTPHTTPInput/output tokens
qwen-tts-2025-04-10Qwen HTTPHTTPInput/output tokens
qwen3-tts-flash-realtimeQwen Realtime TTSrealtime WSSession cumulative characters
qwen3-tts-flash-realtime-2025-11-27Qwen Realtime TTSrealtime WSSession cumulative characters
qwen3-tts-flash-realtime-2025-09-18Qwen Realtime TTSrealtime WSSession cumulative characters
qwen3-tts-instruct-flash-realtimeQwen Realtime TTSrealtime WSSession cumulative characters
qwen3-tts-instruct-flash-realtime-2026-01-22Qwen Realtime TTSrealtime WSSession cumulative characters
qwen3-tts-vc-realtime-2026-01-15Qwen Realtime TTS · Cloningrealtime WSSession cumulative characters
qwen3-tts-vc-realtime-2025-11-27Qwen Realtime TTS · Cloningrealtime WSSession cumulative characters
qwen3-tts-vd-realtime-2026-01-15Qwen Realtime TTS · Designrealtime WSSession cumulative characters
qwen3-tts-vd-realtime-2025-12-16Qwen Realtime TTS · Designrealtime WSSession cumulative characters
qwen-tts-realtimeQwen Realtime TTSrealtime WSTokens per response
qwen-tts-realtime-latestQwen Realtime TTSrealtime WSTokens per response
qwen-tts-realtime-2025-07-15Qwen Realtime TTSrealtime WSTokens per response
MiniMax/speech-2.8-hdMiniMax TTSHTTPCharacters
MiniMax/speech-02-hdMiniMax TTSHTTPCharacters
MiniMax/speech-2.8-turboMiniMax TTSHTTPCharacters
MiniMax/speech-02-turboMiniMax TTSHTTPCharacters
voice-enrollmentVoice customizationcustomization HTTPCreate/manage actions
qwen-voice-enrollmentVoice customizationcustomization HTTPCreate/manage actions
qwen-voice-designVoice customizationcustomization HTTPCreate/manage actions

Choose an API

Use caseRouteOutput
Save complete audio directlyPOST /v1/audio/speechAudio bytes
Native Qwen-Audio / CosyVoicePOST /api/v1/services/audio/tts/SpeechSynthesizerJSON / SSE
Native Qwen HTTP / MiniMaxPOST /api/v1/services/aigc/multimodal-generation/generationJSON / SSE
Duplex Qwen-Audio / CosyVoiceGET /api-ws/v1/inference?model=...Events / binary frames
Qwen realtime TTSGET /api-ws/v1/realtime?model=...Events / Base64 audio
Voice customizationPOST /api/v1/services/audio/tts/customizationJSON

CosyVoice v1 has no native HTTP synthesis route; standard speech bridges to inference WS. Qwen realtime models require WebSocket. AOQ/QUIC transport is outside these HTTP/WS APIs.

Standard Speech

curl --fail-with-body "$MODELSELL_BASE_URL/v1/audio/speech" \
  -H "Authorization: Bearer $MODELSELL_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "cosyvoice-v3-flash",
    "input": "欢迎使用 ModelSell,开始今天的语音创作。",
    "voice": "longanyang",
    "response_format": "mp3",
    "speed": 1.0,
    "extra_body": {"input": {"sample_rate": 24000, "seed": 0}}
  }' --output speech.mp3

Standard fields are model, input (text), voice, response_format, speed, and instructions. Use extra_body.input / extra_body.parameters for native options; explicit zero and false values are preserved. Extensions cannot replace the model or text. stream: true returns binary chunks; stream_format: "sse" preserves native SSE, timestamps and usage. The gateway does not transcode. Qwen HTTP returns WAV synchronously and PCM while streaming.

Usage and Prices

Character models use upstream billing characters; token models use input and output tokens. Qwen3 realtime reports cumulative session characters, while legacy Qwen realtime reports tokens per response. Voice creation and management can have different prices; consult the console rather than assuming queries are free.

Documented models and enabled models

The table identifies API formats; it does not guarantee that all models are enabled in your group. Custom voices must match the target model and provider workspace used when creating them.

Reference: Alibaba Cloud TTS overview. Each guide links to its API reference. Examples contain placeholders, not provider credentials.

On this page