Qwen Realtime TTS
WebSocket events, audio decoding and session usage for 12 Qwen realtime TTS models.
Models: qwen3-tts-flash-realtime, qwen3-tts-flash-realtime-2025-11-27, qwen3-tts-flash-realtime-2025-09-18, qwen3-tts-instruct-flash-realtime, qwen3-tts-instruct-flash-realtime-2026-01-22, qwen3-tts-vc-realtime-2026-01-15, qwen3-tts-vc-realtime-2025-11-27, qwen3-tts-vd-realtime-2026-01-15, qwen3-tts-vd-realtime-2025-12-16, qwen-tts-realtime, qwen-tts-realtime-latest, qwen-tts-realtime-2025-07-15.
This is text-to-speech. For Omni conversations, see Qwen realtime audio protocols.
Connection
GET /api-ws/v1/realtime?model=qwen3-tts-flash-realtime
Authorization: Bearer $MODELSELL_API_KEY
Upgrade: websocketUse wss://api.modelsell.com. The alias /v1/realtime?model=... still expects these Bailian TTS events. Changing the model requires another connection.
Session Configuration
| Session field | Type | Purpose |
|---|---|---|
voice | string | Required; system voice for Flash/Instruct, custom voice for VC/VD |
mode | string | server_commit automatically synthesizes; commit requires client submission |
language_type | string | Target language, e.g. Chinese or Auto |
response_format | string | Qwen3 pcm/wav/mp3/opus; legacy PCM only |
sample_rate | integer | Qwen3 8000/16000/24000/48000; legacy 24000 only |
speech_rate / volume / pitch_rate | number | Qwen3 speed, volume, pitch; unsupported by legacy |
bit_rate | integer | Qwen3 Opus bitrate |
instructions / optimize_instructions | string / boolean | Instruct realtime only |
Each client event needs a unique event_id. Configure the session, wait for session.updated, append text and commit. Decode response.audio.delta.delta as Base64. After response.done, send more text or session.finish, then read through session.finished.
{"event_id":"UNIQUE_EVENT_ID","type":"session.update","session":{"voice":"Cherry","mode":"commit","language_type":"Chinese","response_format":"pcm","sample_rate":24000}}Complete Python Example
Install websocket-client and set the environment variables.
import base64
import json
import os
import uuid
from websocket import create_connection
model = "qwen3-tts-flash-realtime"
base = os.environ["MODELSELL_WS_URL"].rstrip("/")
ws = create_connection(
f"{base}/api-ws/v1/realtime?model={model}",
header=[f"Authorization: Bearer {os.environ['MODELSELL_API_KEY']}"],
timeout=60,
)
def send(kind, **fields):
ws.send(json.dumps({"event_id": str(uuid.uuid4()), "type": kind, **fields}))
try:
send("session.update", session={"voice": "Cherry", "mode": "commit",
"response_format": "pcm", "sample_rate": 24000})
with open("speech.pcm", "wb") as out:
while True:
event = json.loads(ws.recv())
kind = event.get("type")
if kind == "session.updated":
send("input_text_buffer.append", text="欢迎收听今天的节目。")
send("input_text_buffer.commit")
elif kind == "response.audio.delta":
out.write(base64.b64decode(event["delta"]))
elif kind == "response.done":
print("usage:", event.get("response", {}).get("usage", {}))
send("session.finish")
elif kind == "session.finished":
break
elif kind == "error":
raise RuntimeError(event)
finally:
ws.close()The example writes 24 kHz, 16-bit little-endian mono PCM. Configure the player accordingly.
Usage
{"type":"response.done","response":{"id":"RESPONSE_ID","status":"completed","usage":{"characters":18}}}Qwen3 response.usage.characters is cumulative for the session: 9 followed by 18 means 18 total. Legacy Qwen reports input_tokens/output_tokens for each response; deduplicate response IDs and sum them. Read completion events before closing to retain final audio and usage.
References checked 2026-10-06: Client events, Server events, Connection options.