ModelSell 文档
音频与其他百炼语音合成

Qwen 实时 TTS

12 个 Qwen realtime 合成模型的 WebSocket 事件、音频解码与会话用量。

适用型号见 完整模型表:Qwen3 Flash realtime(3 个)、Instruct realtime(2 个)、VC realtime(2 个)、VD realtime(2 个),以及旧 Qwen realtime(3 个)。此接口用于文本转语音;Omni 实时对话见 Qwen 实时音频协议。

模型 ID

qwen3-tts-flash-realtime、qwen3-tts-flash-realtime-2025-11-27、qwen3-tts-flash-realtime-2025-09-18、qwen3-tts-instruct-flash-realtime、qwen3-tts-instruct-flash-realtime-2026-01-22、qwen3-tts-vc-realtime-2026-01-15、qwen3-tts-vc-realtime-2025-11-27、qwen3-tts-vd-realtime-2026-01-15、qwen3-tts-vd-realtime-2025-12-16、qwen-tts-realtime、qwen-tts-realtime-latest、qwen-tts-realtime-2025-07-15。

建立连接

GET /api-ws/v1/realtime?model=qwen3-tts-flash-realtime
Authorization: Bearer $MODELSELL_API_KEY
Upgrade: websocket

使用 wss://api.modelsell.com。兼容路径 /v1/realtime?model=... 仍要求发送下文的百炼 TTS 事件。连接中的模型固定;更换模型需建立新连接。

会话配置

session 字段类型说明
voicestring必填,Flash / Instruct 可用系统音色;VC / VD 使用绑定该模型的专属音色
modestringserver_commit 自动合成;commit 由客户端提交文本
language_typestring目标语种,如 Chinese / Auto
response_formatstringQwen3 可用 pcm / wav / mp3 / opus;旧 Qwen 仅 PCM
sample_rateintegerQwen3 支持 8000 / 16000 / 24000 / 48000;旧 Qwen 仅 24000
speech_rate / volume / pitch_ratenumberQwen3 语速、音量、音调;旧 Qwen 不支持
bit_rateintegerQwen3 Opus 码率,其他格式不使用
instructions / optimize_instructionsstring / boolean仅 Instruct realtime

每个客户端事件的 event_id 都应使用唯一 ID。

{"event_id":"UNIQUE_EVENT_ID","type":"session.update","session":{"voice":"Cherry","mode":"commit","language_type":"Chinese","response_format":"pcm","sample_rate":24000}}

收到 session.updated 后追加文本,再发送 input_text_buffer.commit。读取 response.audio.delta,其 delta 为 Base64 音频;response.done 表示本次回复完成。可继续追加下一句,最终发送 session.finish,读取剩余事件直到 session.finished。

Python 完整示例

安装 websocket-client,先按 概览 设置环境变量:

import base64
import json
import os
import uuid
from websocket import create_connection

model = "qwen3-tts-flash-realtime"
base = os.environ["MODELSELL_WS_URL"].rstrip("/")
ws = create_connection(
    f"{base}/api-ws/v1/realtime?model={model}",
    header=[f"Authorization: Bearer {os.environ['MODELSELL_API_KEY']}"],
    timeout=60,
)

def send(kind, **fields):
    ws.send(json.dumps({"event_id": str(uuid.uuid4()), "type": kind, **fields}))

try:
    send("session.update", session={"voice": "Cherry", "mode": "commit",
         "response_format": "pcm", "sample_rate": 24000})
    with open("speech.pcm", "wb") as out:
        while True:
            event = json.loads(ws.recv())
            kind = event.get("type")
            if kind == "session.updated":
                send("input_text_buffer.append", text="欢迎收听今天的节目。")
                send("input_text_buffer.commit")
            elif kind == "response.audio.delta":
                out.write(base64.b64decode(event["delta"]))
            elif kind == "response.done":
                print("usage:", event.get("response", {}).get("usage", {}))
                send("session.finish")
            elif kind == "session.finished":
                break
            elif kind == "error":
                raise RuntimeError(event)
finally:
    ws.close()

示例 PCM 为 24 kHz、16 位小端、单声道;播放器需设置相同参数。

返回与收费

{"type":"response.done","response":{"id":"RESPONSE_ID","status":"completed","usage":{"characters":18}}}

Qwen3 的 response.usage.characters 是会话累计字符数,例如 9、18 应最终按 18 计量。旧 Qwen 的各 response 分别报告 input_tokens / output_tokens,按 response ID 去重后相加。关闭客户端前读取完成事件,避免丢失尾部音频和用量。

官方参考:客户端事件、服务端事件、接入方式。核对日期:2026-10-06。

On this page