Voice Cloning, Design and Management
Create and manage voices with voice-enrollment, qwen-voice-enrollment and qwen-voice-design.
Route and Services
POST /api/v1/services/audio/tts/customization
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json| Service model | Purpose | Create action | Voice output |
|---|---|---|---|
| voice-enrollment | Qwen-Audio/CosyVoice cloning and supported voice design | create_voice | output.voice_id |
| qwen-voice-enrollment | Qwen3 VC cloning | create | output.voice |
| qwen-voice-design | Qwen3 VD design | create | output.voice |
input.target_model must match subsequent synthesis. Both the customization service and synthesis model require token permission. Voices must match the region and workspace. MiniMax cloning uses a different route.
Qwen-Audio / CosyVoice Cloning
curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/audio/tts/customization" \
-H "Authorization: Bearer $MODELSELL_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"voice-enrollment","input":{"action":"create_voice","target_model":"cosyvoice-v3-flash","prefix":"studio","url":"https://example.com/my-voice-sample.wav","language_hints":["zh"]}}'Use a real publicly accessible URL. prefix is at most 10 alphanumeric characters. Optional enable_preprocess/max_prompt_audio_length depend on the target. enable_volume_normalization is the string "true"/"false" for this API.
Qwen3 VC Cloning
{"model":"qwen-voice-enrollment","input":{"action":"create","target_model":"qwen3-tts-vc-2026-01-22","preferred_name":"studio_reader","audio":{"data":"https://example.com/my-voice-sample.wav"},"language":"zh"}}audio.data accepts a URL or data URI, e.g. data:audio/wav;base64,BASE64_SAMPLE. Optional text supplies the sample transcript. preferred_name allows up to 16 letters, digits and underscores. Check fallback_mode/fallback_reason for degraded creation.
Qwen3 VD Design
{
"model": "qwen-voice-design",
"input": {
"action": "create", "target_model": "qwen3-tts-vd-2026-01-26",
"preferred_name": "studio_reader",
"voice_prompt": "成年女性,声音清晰平稳,语速适中,适合讲述科普故事。",
"preview_text": "欢迎收听今天的科学故事,让我们一起了解生活中的新发现。"
},
"parameters": {"sample_rate": 24000, "response_format": "wav"}
}voice_prompt describes the voice; preview_text is the audition text. For a supported CosyVoice design model, use voice-enrollment/create_voice and a compatible target such as cosyvoice-v3.5-plus; other requirements follow its reference.
Response and Synthesis
{"request_id":"REQUEST_ID","output":{"voice":"CUSTOM_VOICE","target_model":"qwen3-tts-vd-2026-01-26","preview_audio":{"data":"BASE64_WAV","sample_rate":24000}},"usage":{"count":1}}Qwen-Audio/CosyVoice returns output.voice_id. Preview audio may appear at output.preview_audio or other documented fields; the full JSON is preserved. Use the ID with Qwen HTTP, Realtime TTS, or Qwen-Audio/CosyVoice, matching the creation target.
Management Actions
All actions are POSTs to the same customization route.
| Service | Action | Additional input |
|---|---|---|
| voice-enrollment | list_voice | prefix, page_index, page_size |
| voice-enrollment | query_voice | voice_id |
| voice-enrollment | update_voice | voice_id, url |
| voice-enrollment | delete_voice | voice_id |
| qwen-voice-enrollment | list / delete | Pagination / voice |
| qwen-voice-design | list / delete | Pagination / voice; manages designed voices |
{"model":"voice-enrollment","input":{"action":"query_voice","voice_id":"CUSTOM_VOICE_ID"}}{"model":"qwen-voice-design","input":{"action":"delete","voice":"CUSTOM_DESIGNED_VOICE"}}Qwen does not support CosyVoice query_voice/update_voice. MiniMax does not support these management operations. Create/update/query can have different billing rules; see usage.
References checked 2026-10-06: Cloning and management API, Voice design.