ModelSell Docs
Audio & MoreBailian Speech Synthesis

MiniMax Speech Synthesis and Cloning

HTTP, SSE, hex audio, subtitles and voice cloning for 4 MiniMax Speech models.

Models: MiniMax/speech-2.8-hd, MiniMax/speech-02-hd, MiniMax/speech-2.8-turbo, MiniMax/speech-02-turbo.

Route and Parameters

POST /api/v1/services/aigc/multimodal-generation/generation
Authorization: Bearer $MODELSELL_API_KEY
Content-Type: application/json
FieldTypePurpose
model / input.textstringRequired model and text
input.voice_setting.voice_idstringRequired system/custom voice
input.voice_setting.speed / vol / pitch / emotionnumber / stringOptional voice controls
input.audio_settingobjectformat, sample_rate, bitrate, channel
input.output_formatstringhex (default) or url; streaming uses hex
input.stream_options.exclude_aggregated_audiobooleanSet true to suppress full-audio tail frame
input.pronunciation_dict / timbre_weights / voice_modifyobject / array / objectPronunciation, voice mixing, effects
input.language_boost / latex_read / text_normalizationstring / boolean / booleanLanguage and text options
input.subtitle_enable / aigc_watermarkbooleanSynchronous subtitles / watermark
curl --fail-with-body "$MODELSELL_BASE_URL/api/v1/services/aigc/multimodal-generation/generation" \
  -H "Authorization: Bearer $MODELSELL_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"MiniMax/speech-2.8-turbo","input":{"text":"欢迎收听我们的创作分享。","voice_setting":{"voice_id":"female-shaonv","speed":1.0,"pitch":0},"audio_setting":{"format":"mp3","sample_rate":24000},"output_format":"hex"}}' \
  --output response.json

Decode Output

{"request_id":"REQUEST_ID","output":{"data":{"audio":"HEX_AUDIO","status":2},"extra_info":{"usage_characters":12},"base_resp":{"status_code":0,"status_msg":""}},"usage":{"characters":12}}

Decode hex, not Base64:

import json
from pathlib import Path

result = json.loads(Path("response.json").read_text())
Path("speech.mp3").write_bytes(bytes.fromhex(result["output"]["data"]["audio"]))

With output_format url, output.data.audio is a temporary download URL. Subtitles may appear at output.data.subtitle_file. Nonzero output.base_resp.status_code indicates a provider error.

SSE and Standard Speech

Add X-DashScope-SSE: enable for hex chunks. With input.stream_options.exclude_aggregated_audio true, decode and append chunks in order without duplicating the complete-audio tail frame. Standard speech decodes audio directly. voice maps to voice_setting.voice_id; speed to voice_setting.speed; response_format to audio_setting.format. Put other options in extra_body.input.

Voice Cloning

Use the same multigen route with input.action voice_clone, not customization:

{"model":"MiniMax/speech-2.8-turbo","input":{"action":"voice_clone","voice_id":"MyStudioVoice20261006","audio_url":"https://example.com/my-voice-sample.wav","text":"这是一段用于确认新音色的试听文本。","need_noise_reduction":false,"need_volume_normalization":false}}

Supply a real publicly accessible sample URL and a unique voice_id. Check output.base_resp and preview output.demo_audio, then synthesize using that voice_id. Bailian does not expose MiniMax list/query/update/delete voice actions. Preview synthesis and first use may incur provider charges; ModelSell billing follows the console rules.

References checked 2026-10-06: Synthesis API, Cloning API.

On this page