声音复刻、设计与音色管理
voice-enrollment、qwen-voice-enrollment 与 qwen-voice-design 的创建、返回和管理操作。
统一入口与服务选择
POST /api/v1/services/audio/tts/customization
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json| 服务 model | 用途 | 创建 action | 音色返回字段 |
|---|---|---|---|
voice-enrollment | Qwen-Audio / CosyVoice 复刻,以及受支持型号的声音设计 | create_voice | output.voice_id |
qwen-voice-enrollment | Qwen3 VC 复刻 | create | output.voice |
qwen-voice-design | Qwen3 VD 设计 | create | output.voice |
input.target_model 是后续合成型号;创建音色的服务也必须在 Token 可用模型中。区域、业务空间和目标模型要与合成请求匹配。MiniMax 使用 另一条克隆入口。
Qwen-Audio / CosyVoice 复刻
curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/audio/tts/customization" \
-H "Authorization: Bearer $MODELSELL_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"voice-enrollment","input":{"action":"create_voice","target_model":"cosyvoice-v3-flash","prefix":"studio","url":"https://example.com/my-voice-sample.wav","language_hints":["zh"]}}'样本 url 必須公网可访问。prefix 用英文字母或数字,最长 10 个字符。可选 enable_preprocess、max_prompt_audio_length 依目标模型而定;enable_volume_normalization 按此接口要求传字符串 "true" / "false"。
Qwen3 VC 复刻
{"model":"qwen-voice-enrollment","input":{"action":"create","target_model":"qwen3-tts-vc-2026-01-22","preferred_name":"studio_reader","audio":{"data":"https://example.com/my-voice-sample.wav"},"language":"zh"}}audio.data 接受 URL 或音频 data URI,如 data:audio/wav;base64,BASE64_SAMPLE。可选 text 提供样本转写。preferred_name 使用字母、数字或下划线,最长 16 个字符。检查 fallback_mode / fallback_reason,判断复刻是否降级。
Qwen3 VD 设计
{
"model": "qwen-voice-design",
"input": {
"action": "create", "target_model": "qwen3-tts-vd-2026-01-26",
"preferred_name": "studio_reader",
"voice_prompt": "成年女性,声音清晰平稳,语速适中,适合讲述科普故事。",
"preview_text": "欢迎收听今天的科学故事,让我们一起了解生活中的新发现。"
},
"parameters": {"sample_rate": 24000, "response_format": "wav"}
}voice_prompt 描述声音,preview_text 指定试听文本;CosyVoice 设计音色时,将服务改为 voice-enrollment、action 改为 create_voice,并选支持设计的目标模型(例如 cosyvoice-v3.5-plus)。其他必填字段以该模型官方参考为准。
创建响应与后续合成
{"request_id":"REQUEST_ID","output":{"voice":"CUSTOM_VOICE","target_model":"qwen3-tts-vd-2026-01-26","preview_audio":{"data":"BASE64_WAV","sample_rate":24000}},"usage":{"count":1}}Qwen-Audio / CosyVoice 从 output.voice_id 取值。试听可在 output.preview_audio 或其他官方输出字段中返回;网关保留完整 JSON。将新音色用于 Qwen HTTP、实时 TTS 或 Qwen-Audio / CosyVoice,不要将同一音色随意用于另一型号。
查询、列表、更新与删除
所有操作均为同一路径的 POST,区别在 JSON:
| 服务 | action | input 额外字段 | 结果 / 限制 |
|---|---|---|---|
voice-enrollment | list_voice | prefix、page_index、page_size | output.voice_list |
voice-enrollment | query_voice | voice_id | 状态、目标模型、样本链接 |
voice-enrollment | update_voice | voice_id、url | 更新复刻样本 |
voice-enrollment | delete_voice | voice_id | 删除对应音色 |
qwen-voice-enrollment | list | page_index、page_size | 复刻音色列表 |
qwen-voice-enrollment | delete | voice | 删除复刻音色 |
qwen-voice-design | list / delete | 分页字段 / voice | 设计音色由设计服务管理 |
{"model":"voice-enrollment","input":{"action":"query_voice","voice_id":"CUSTOM_VOICE_ID"}}{"model":"qwen-voice-design","input":{"action":"delete","voice":"CUSTOM_DESIGNED_VOICE"}}Qwen 系列不支持 CosyVoice 的 query_voice / update_voice;MiniMax 不支持上述管理操作。创建、更新与查询的价格规则可能不同,详见 用量说明。