Overview
Paper: FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech
Authors: Puneet Mathur, Dinesh Manocha
Subjects: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
arXiv: 2609.35791 [cs.CL] — Submitted 16 Sep 2026
DOI: 10.48550/arXiv.2609.35791
Abstract
Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection (VAD) lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages.
We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference.
We further introduce confidence-gated endpoint commitment to control interruption versus delay, and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on the TurnBench dev set — 0.853 (at FP ≤ 0.10) in a zero-shot setting.
These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.
Key Contributions
- Semantic endpoint detection as causal audio-language reasoning: Rather than relying on acoustic cues alone or on cascaded ASR pipelines, FD-VAD treats the decision of whether a speaker has finished their turn as a direct reasoning problem over streaming audio.
- ASR-free streaming architecture: FD-VAD pairs a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model. This design avoids the transcription dependence and latency overhead of cascaded ASR-based endpointing.
- Last-chunk training objective: The model is trained to make streaming inference decisions from bounded causal audio windows, aligning training with real-time deployment conditions.
- Confidence-gated endpoint commitment: A mechanism that balances the trade-off between premature interruption and excessive delay by gating endpoint decisions on model confidence.
- Boundary-focused hard-negative sampling: A training strategy that improves discrimination around ambiguous turn boundaries — precisely the cases where false endpointing is most costly.
Results
Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers. On the TurnBench dev set, it achieves the highest EOT (end-of-turn) recall among qualifying systems — 0.853 at FP ≤ 0.10 — in a zero-shot setting.
Significance
These results demonstrate that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking. This has practical implications for full-duplex voice interfaces, where low-latency, semantically aware turn-taking remains a central challenge. By removing the transcription stage, FD-VAD reduces pipeline complexity and latency while preserving the semantic understanding needed to distinguish hesitation from completed intent.
Metadata
- Cite as: arXiv:2609.35791 [cs.CL]
- Version: v1 (16 Sep 2026)
- Submission history: [v1] Wed, 16 Sep 2026 20:35:46 UTC (2,522 KB)
via ArXiv CL+LG
