Bailian TTS Models and Integration
API, format, voice and billing map for 36 synthesis models and 3 voice services.
These guides cover Qwen-Audio-TTS, CosyVoice, Qwen-TTS and MiniMax. The documented inventory was checked on 2026-10-06. Model availability, token permissions and current prices depend on your ModelSell group.
Authentication
Use a ModelSell API key and the ModelSell base URL. The channel manages provider credentials and workspace addresses.
export MODELSELL_BASE_URL="https://api.modelsell.com"
export MODELSELL_WS_URL="wss://api.modelsell.com"
export MODELSELL_API_KEY="YOUR_MODELSELL_API_KEY"HTTP uses Authorization: Bearer $MODELSELL_API_KEY and Content-Type: application/json. WebSocket handshakes use the same Authorization header.
Complete Model and Documentation Map
| Model ID | Integration guide | Protocol | Usage unit |
|---|---|---|---|
qwen-audio-3.0-tts-plus | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
qwen-audio-3.1-tts-flash | Qwen-Audio / CosyVoice | HTTP / inference WS | Input/output tokens |
qwen-audio-3.0-tts-flash | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
cosyvoice-v3.5-plus | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
cosyvoice-v3.5-flash | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
cosyvoice-v3-plus | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
cosyvoice-v3-flash | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
cosyvoice-v2 | Qwen-Audio / CosyVoice | HTTP / inference WS | Characters |
cosyvoice-v1 | Qwen-Audio / CosyVoice | inference WS;speech bridge | Characters |
qwen3-tts-flash | Qwen HTTP | HTTP | Characters |
qwen3-tts-flash-2025-11-27 | Qwen HTTP | HTTP | Characters |
qwen3-tts-flash-2025-09-18 | Qwen HTTP | HTTP | Characters |
qwen3-tts-instruct-flash | Qwen HTTP | HTTP | Characters |
qwen3-tts-instruct-flash-2026-01-26 | Qwen HTTP | HTTP | Characters |
qwen3-tts-vc-2026-01-22 | Qwen HTTP · Cloning | HTTP | Characters |
qwen3-tts-vd-2026-01-26 | Qwen HTTP · Design | HTTP | Characters |
qwen-tts | Qwen HTTP | HTTP | Input/output tokens |
qwen-tts-latest | Qwen HTTP | HTTP | Input/output tokens |
qwen-tts-2025-05-22 | Qwen HTTP | HTTP | Input/output tokens |
qwen-tts-2025-04-10 | Qwen HTTP | HTTP | Input/output tokens |
qwen3-tts-flash-realtime | Qwen Realtime TTS | realtime WS | Session cumulative characters |
qwen3-tts-flash-realtime-2025-11-27 | Qwen Realtime TTS | realtime WS | Session cumulative characters |
qwen3-tts-flash-realtime-2025-09-18 | Qwen Realtime TTS | realtime WS | Session cumulative characters |
qwen3-tts-instruct-flash-realtime | Qwen Realtime TTS | realtime WS | Session cumulative characters |
qwen3-tts-instruct-flash-realtime-2026-01-22 | Qwen Realtime TTS | realtime WS | Session cumulative characters |
qwen3-tts-vc-realtime-2026-01-15 | Qwen Realtime TTS · Cloning | realtime WS | Session cumulative characters |
qwen3-tts-vc-realtime-2025-11-27 | Qwen Realtime TTS · Cloning | realtime WS | Session cumulative characters |
qwen3-tts-vd-realtime-2026-01-15 | Qwen Realtime TTS · Design | realtime WS | Session cumulative characters |
qwen3-tts-vd-realtime-2025-12-16 | Qwen Realtime TTS · Design | realtime WS | Session cumulative characters |
qwen-tts-realtime | Qwen Realtime TTS | realtime WS | Tokens per response |
qwen-tts-realtime-latest | Qwen Realtime TTS | realtime WS | Tokens per response |
qwen-tts-realtime-2025-07-15 | Qwen Realtime TTS | realtime WS | Tokens per response |
MiniMax/speech-2.8-hd | MiniMax TTS | HTTP | Characters |
MiniMax/speech-02-hd | MiniMax TTS | HTTP | Characters |
MiniMax/speech-2.8-turbo | MiniMax TTS | HTTP | Characters |
MiniMax/speech-02-turbo | MiniMax TTS | HTTP | Characters |
voice-enrollment | Voice customization | customization HTTP | Create/manage actions |
qwen-voice-enrollment | Voice customization | customization HTTP | Create/manage actions |
qwen-voice-design | Voice customization | customization HTTP | Create/manage actions |
Choose an API
| Use case | Route | Output |
|---|---|---|
| Save complete audio directly | POST /v1/audio/speech | Audio bytes |
| Native Qwen-Audio / CosyVoice | POST /api/v1/services/audio/tts/SpeechSynthesizer | JSON / SSE |
| Native Qwen HTTP / MiniMax | POST /api/v1/services/aigc/multimodal-generation/generation | JSON / SSE |
| Duplex Qwen-Audio / CosyVoice | GET /api-ws/v1/inference?model=... | Events / binary frames |
| Qwen realtime TTS | GET /api-ws/v1/realtime?model=... | Events / Base64 audio |
| Voice customization | POST /api/v1/services/audio/tts/customization | JSON |
CosyVoice v1 has no native HTTP synthesis route; standard speech bridges to inference WS. Qwen realtime models require WebSocket. AOQ/QUIC transport is outside these HTTP/WS APIs.
Standard Speech
curl --fail-with-body "$MODELSELL_BASE_URL/v1/audio/speech" \
-H "Authorization: Bearer $MODELSELL_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "cosyvoice-v3-flash",
"input": "欢迎使用 ModelSell,开始今天的语音创作。",
"voice": "longanyang",
"response_format": "mp3",
"speed": 1.0,
"extra_body": {"input": {"sample_rate": 24000, "seed": 0}}
}' --output speech.mp3Standard fields are model, input (text), voice, response_format, speed, and instructions. Use extra_body.input / extra_body.parameters for native options; explicit zero and false values are preserved. Extensions cannot replace the model or text. stream: true returns binary chunks; stream_format: "sse" preserves native SSE, timestamps and usage. The gateway does not transcode. Qwen HTTP returns WAV synchronously and PCM while streaming.
Usage and Prices
Character models use upstream billing characters; token models use input and output tokens. Qwen3 realtime reports cumulative session characters, while legacy Qwen realtime reports tokens per response. Voice creation and management can have different prices; consult the console rather than assuming queries are free.
Documented models and enabled models
The table identifies API formats; it does not guarantee that all models are enabled in your group. Custom voices must match the target model and provider workspace used when creating them.
Reference: Alibaba Cloud TTS overview. Each guide links to its API reference. Examples contain placeholders, not provider credentials.