ModelSell Docs
Audio & MoreBailian Speech Synthesis

Voice Cloning, Design and Management

Create and manage voices with voice-enrollment, qwen-voice-enrollment and qwen-voice-design.

Route and Services

POST /api/v1/services/audio/tts/customization
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json
Service modelPurposeCreate actionVoice output
voice-enrollmentQwen-Audio/CosyVoice cloning and supported voice designcreate_voiceoutput.voice_id
qwen-voice-enrollmentQwen3 VC cloningcreateoutput.voice
qwen-voice-designQwen3 VD designcreateoutput.voice

input.target_model must match subsequent synthesis. Both the customization service and synthesis model require token permission. Voices must match the region and workspace. MiniMax cloning uses a different route.

Qwen-Audio / CosyVoice Cloning

curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/audio/tts/customization" \
  -H "Authorization: Bearer $MODELSELL_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"voice-enrollment","input":{"action":"create_voice","target_model":"cosyvoice-v3-flash","prefix":"studio","url":"https://example.com/my-voice-sample.wav","language_hints":["zh"]}}'

Use a real publicly accessible URL. prefix is at most 10 alphanumeric characters. Optional enable_preprocess/max_prompt_audio_length depend on the target. enable_volume_normalization is the string "true"/"false" for this API.

Qwen3 VC Cloning

{"model":"qwen-voice-enrollment","input":{"action":"create","target_model":"qwen3-tts-vc-2026-01-22","preferred_name":"studio_reader","audio":{"data":"https://example.com/my-voice-sample.wav"},"language":"zh"}}

audio.data accepts a URL or data URI, e.g. data:audio/wav;base64,BASE64_SAMPLE. Optional text supplies the sample transcript. preferred_name allows up to 16 letters, digits and underscores. Check fallback_mode/fallback_reason for degraded creation.

Qwen3 VD Design

{
  "model": "qwen-voice-design",
  "input": {
    "action": "create", "target_model": "qwen3-tts-vd-2026-01-26",
    "preferred_name": "studio_reader",
    "voice_prompt": "成年女性,声音清晰平稳,语速适中,适合讲述科普故事。",
    "preview_text": "欢迎收听今天的科学故事,让我们一起了解生活中的新发现。"
  },
  "parameters": {"sample_rate": 24000, "response_format": "wav"}
}

voice_prompt describes the voice; preview_text is the audition text. For a supported CosyVoice design model, use voice-enrollment/create_voice and a compatible target such as cosyvoice-v3.5-plus; other requirements follow its reference.

Response and Synthesis

{"request_id":"REQUEST_ID","output":{"voice":"CUSTOM_VOICE","target_model":"qwen3-tts-vd-2026-01-26","preview_audio":{"data":"BASE64_WAV","sample_rate":24000}},"usage":{"count":1}}

Qwen-Audio/CosyVoice returns output.voice_id. Preview audio may appear at output.preview_audio or other documented fields; the full JSON is preserved. Use the ID with Qwen HTTP, Realtime TTS, or Qwen-Audio/CosyVoice, matching the creation target.

Management Actions

All actions are POSTs to the same customization route.

ServiceActionAdditional input
voice-enrollmentlist_voiceprefix, page_index, page_size
voice-enrollmentquery_voicevoice_id
voice-enrollmentupdate_voicevoice_id, url
voice-enrollmentdelete_voicevoice_id
qwen-voice-enrollmentlist / deletePagination / voice
qwen-voice-designlist / deletePagination / voice; manages designed voices
{"model":"voice-enrollment","input":{"action":"query_voice","voice_id":"CUSTOM_VOICE_ID"}}
{"model":"qwen-voice-design","input":{"action":"delete","voice":"CUSTOM_DESIGNED_VOICE"}}

Qwen does not support CosyVoice query_voice/update_voice. MiniMax does not support these management operations. Create/update/query can have different billing rules; see usage.

References checked 2026-10-06: Cloning and management API, Voice design.

On this page