Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

Alibaba’s Qwen team has released Qwen-Audio-3.1, a five-model audio stack spanning ASR, TTS, and realtime interaction. The flagship model is Qwen-Audio-3.1-Realtime, a full-duplex speech model built for voice agents that can call tools. Qwen also cut prices: roughly 85% on Realtime, about 70% on TTS, and up to 95% on ASR.

Is It Deployable?

Yes, as a managed API. qwen-audio-3.1-realtime-plus is live on QwenCloud over WebSocket. No open weights were announced.

What Ships on QwenCloud

The model page lists text and audio as both input and output. Context is 262K tokens, with 245K max input and 16K max output. Default limits are 60 requests and 100K tokens per minute. Pricing is $6.4 per 1M audio input tokens and $0.8 per 1M text input tokens. Text and audio output costs $24 per 1M tokens, with output text not charged. Key features include function calling, web search, structured outputs, context cache, and fine-tuning.

A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline long-audio transcription. It supports hot words, speaker separation, punctuation, and multilingual plus Chinese dialect recognition. It costs $0.15 input and $0.47 output per 1M tokens.

Architecture: Two Models Behind One Voice

The system runs two models with the same Audio Encoder and LLM design. A full-duplex decision model predicts whether to keep listening, speak, or stop—learning when to yield the turn in realtime conversation. The interaction itself is handled by a generative model, giving the stack a unified audio understanding and generation backbone.

In 2026, this kind of full-duplex architecture is becoming the default for voice agents: instead of alternating turns, the model continuously tracks the audio stream and decides when and what to speak, which is critical for natural, interruptible dialogue with tool use.

What This Means for Voice Agents

Qwen-Audio-3.1-Realtime brings three capabilities together in one managed endpoint: realtime speech understanding, realtime speech generation, and function calling. That combination makes it viable for production voice agents that need to reason, act, and respond in a single session. With the steep price cuts across ASR, TTS, and realtime audio, Alibaba is positioning QwenCloud as a cost-competitive platform for always-on audio workloads in 2026.

via MarkTechPost

Related