ModelSell Docs
Audio & MoreBailian Speech Synthesis

Qwen Realtime TTS

WebSocket events, audio decoding and session usage for 12 Qwen realtime TTS models.

Models: qwen3-tts-flash-realtime, qwen3-tts-flash-realtime-2025-11-27, qwen3-tts-flash-realtime-2025-09-18, qwen3-tts-instruct-flash-realtime, qwen3-tts-instruct-flash-realtime-2026-01-22, qwen3-tts-vc-realtime-2026-01-15, qwen3-tts-vc-realtime-2025-11-27, qwen3-tts-vd-realtime-2026-01-15, qwen3-tts-vd-realtime-2025-12-16, qwen-tts-realtime, qwen-tts-realtime-latest, qwen-tts-realtime-2025-07-15.

This is text-to-speech. For Omni conversations, see Qwen realtime audio protocols.

Connection

GET /api-ws/v1/realtime?model=qwen3-tts-flash-realtime
Authorization: Bearer $MODELSELL_API_KEY
Upgrade: websocket

Use wss://api.modelsell.com. The alias /v1/realtime?model=... still expects these Bailian TTS events. Changing the model requires another connection.

Session Configuration

Session fieldTypePurpose
voicestringRequired; system voice for Flash/Instruct, custom voice for VC/VD
modestringserver_commit automatically synthesizes; commit requires client submission
language_typestringTarget language, e.g. Chinese or Auto
response_formatstringQwen3 pcm/wav/mp3/opus; legacy PCM only
sample_rateintegerQwen3 8000/16000/24000/48000; legacy 24000 only
speech_rate / volume / pitch_ratenumberQwen3 speed, volume, pitch; unsupported by legacy
bit_rateintegerQwen3 Opus bitrate
instructions / optimize_instructionsstring / booleanInstruct realtime only

Each client event needs a unique event_id. Configure the session, wait for session.updated, append text and commit. Decode response.audio.delta.delta as Base64. After response.done, send more text or session.finish, then read through session.finished.

{"event_id":"UNIQUE_EVENT_ID","type":"session.update","session":{"voice":"Cherry","mode":"commit","language_type":"Chinese","response_format":"pcm","sample_rate":24000}}

Complete Python Example

Install websocket-client and set the environment variables.

import base64
import json
import os
import uuid
from websocket import create_connection

model = "qwen3-tts-flash-realtime"
base = os.environ["MODELSELL_WS_URL"].rstrip("/")
ws = create_connection(
    f"{base}/api-ws/v1/realtime?model={model}",
    header=[f"Authorization: Bearer {os.environ['MODELSELL_API_KEY']}"],
    timeout=60,
)

def send(kind, **fields):
    ws.send(json.dumps({"event_id": str(uuid.uuid4()), "type": kind, **fields}))

try:
    send("session.update", session={"voice": "Cherry", "mode": "commit",
         "response_format": "pcm", "sample_rate": 24000})
    with open("speech.pcm", "wb") as out:
        while True:
            event = json.loads(ws.recv())
            kind = event.get("type")
            if kind == "session.updated":
                send("input_text_buffer.append", text="欢迎收听今天的节目。")
                send("input_text_buffer.commit")
            elif kind == "response.audio.delta":
                out.write(base64.b64decode(event["delta"]))
            elif kind == "response.done":
                print("usage:", event.get("response", {}).get("usage", {}))
                send("session.finish")
            elif kind == "session.finished":
                break
            elif kind == "error":
                raise RuntimeError(event)
finally:
    ws.close()

The example writes 24 kHz, 16-bit little-endian mono PCM. Configure the player accordingly.

Usage

{"type":"response.done","response":{"id":"RESPONSE_ID","status":"completed","usage":{"characters":18}}}

Qwen3 response.usage.characters is cumulative for the session: 9 followed by 18 means 18 total. Legacy Qwen reports input_tokens/output_tokens for each response; deduplicate response IDs and sum them. Read completion events before closing to retain final audio and usage.

References checked 2026-10-06: Client events, Server events, Connection options.

On this page