FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

Overview


Paper: FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech


Authors: Puneet Mathur, Dinesh Manocha


Subjects: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)


arXiv: 2609.35791 [cs.CL] — Submitted 16 Sep 2026


DOI: 10.48550/arXiv.2609.35791




Abstract


Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection (VAD) lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages.


We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference.


We further introduce confidence-gated endpoint commitment to control interruption versus delay, and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on the TurnBench dev set — 0.853 (at FP ≤ 0.10) in a zero-shot setting.


These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.




Key Contributions


  • Semantic endpoint detection as causal audio-language reasoning: Rather than relying on acoustic cues alone or on cascaded ASR pipelines, FD-VAD treats the decision of whether a speaker has finished their turn as a direct reasoning problem over streaming audio.

  • ASR-free streaming architecture: FD-VAD pairs a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model. This design avoids the transcription dependence and latency overhead of cascaded ASR-based endpointing.

  • Last-chunk training objective: The model is trained to make streaming inference decisions from bounded causal audio windows, aligning training with real-time deployment conditions.

  • Confidence-gated endpoint commitment: A mechanism that balances the trade-off between premature interruption and excessive delay by gating endpoint decisions on model confidence.

  • Boundary-focused hard-negative sampling: A training strategy that improves discrimination around ambiguous turn boundaries — precisely the cases where false endpointing is most costly.



Results


Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers. On the TurnBench dev set, it achieves the highest EOT (end-of-turn) recall among qualifying systems — 0.853 at FP ≤ 0.10 — in a zero-shot setting.




Significance


These results demonstrate that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking. This has practical implications for full-duplex voice interfaces, where low-latency, semantically aware turn-taking remains a central challenge. By removing the transcription stage, FD-VAD reduces pipeline complexity and latency while preserving the semantic understanding needed to distinguish hesitation from completed intent.




Metadata


  • Cite as: arXiv:2609.35791 [cs.CL]
  • Version: v1 (16 Sep 2026)
  • Submission history: [v1] Wed, 16 Sep 2026 20:35:46 UTC (2,522 KB)

via ArXiv CL+LG

Related