PolyAI's Dialog-RSN-1: An Audio-Native Dialog Model Combining Turn-Taking, Speech Recognition, Function

audio-native dialog modeldialog-rsn-1enterprise aifunction callingpolyairesponse generationspeech recognitionturn-takingvoice ai

Overview

PolyAI has released Dialog-RSN-1, a dialog model that processes the caller's audio directly rather than relying on transcripts. It integrates turn-taking, speech recognition, function calling, and response generation into a single audio-native model, and it's currently managing live production calls.

Key Takeaways

  • Dialog-RSN-1 is audio-aware on the input side only; TTS remains separate, ensuring the output voice stays controllable.
  • It operates as a request-based LLM probed on demand, not an always-on stream that ties up a GPU.
  • Turn-taking is the model's first output token: EMPTY, ONGOING, or COMPLETE.
  • PolyAI reports sub-300ms response times, a +11% relative containment improvement at a restaurant group, and a −37% latency reduction at an insurer.
  • English-only at launch, delivered via PolyAI's platform rather than open weights or a public API.

Deployment and Target Users

Dialog-RSN-1 is deployable, but only through PolyAI: there are no open weights or public API yet. Existing customers can enable it immediately, while new customers can request early access.

2026 Context: The Shift Toward Audio-Native AI

Dialog-RSN-1 arrives amid a broader industry push toward audio-native AI systems. In 2026, enterprises are moving beyond cascaded speech-to-text pipelines toward models that directly interpret acoustic signals—reducing latency and improving natural conversational flow. PolyAI's emphasis on request-based inference (rather than persistent streaming) aligns with cost-efficiency trends, as companies seek GPU optimization without sacrificing responsiveness. This model positions PolyAI to compete with other voice AI providers like OpenAI's real-time API and Google's Gemini Live, but with a sharper focus on enterprise telephony use cases.

As PolyAI continues to expand its footprint, the company's ability to deliver low-latency, turn-taking-aware dialog models may set a benchmark for 2026 enterprise voice AI deployments.

via MarkTechPost

Related