CompleteDate: Sep 8, 2026

Experiments

No experiments published for this milestone yet.

Selective User-Stream Routing for Full-Duplex Spoken Dialogue

Motivation

Full-duplex spoken dialogue models continuously receive user speech while generating their own response. However, not all incoming speech should influence generation equally. Interfering speakers, uncertain speech, backchannels, and genuine interruptions differ in both reliability and relevance.

Recent work addresses parts of this problem. IRAF1 uses reliability-aware gating to suppress unreliable user representations before fusion, while Lu et al.2 show that channel fusion provides stronger semantic grounding but cross-attention offers greater robustness to disruptive input. These findings suggest that user-stream integration should be adaptive rather than fixed.

Gap

Current methods largely treat selection and routing separately. Reliability-aware approaches adapt how much user information is passed forward but use a fixed integration mechanism, while routing studies compare fixed fusion strategies without adapting the route to the incoming speech.

This raises a broader question: should a full-duplex model jointly decide what user information should influence generation, how strongly, and through which pathway?

Research Questions

Main RQ: How can full-duplex spoken dialogue models selectively route incoming speech based on both its reliability and its relevance to the ongoing interaction, rather than treating all speech sources (target speaker/speakers, environmental noises/informative cues, etc.) as equally influential?

SubRQ1: Can user-stream routing dynamically adapt to the reliability of incoming speech? For example, highly reliable speech may benefit from channel fusion and strong semantic grounding, while moderately reliable speech may be better accessed through cross-attention to preserve generation robustness.

SubRQ2: Since routing strongly depends on reliability, should selection and routing of incoming user information be learned jointly rather than treated as separate architectural decisions?

SubRQ3: What information should determine how incoming user speech influences ongoing generation? For example, interruptions and backchannels may require different routing behaviours even when both come from the target speaker.

These questions may further connect to speech representation learning and context-aware dialogue modelling.

Possible Approach

I would initially focus on SubRQ1 by developing a reliability-aware dynamic routing mechanism. Rather than using reliability only as a scalar gate that uniformly scales the user representation, the model would learn a richer reliability representation that distinguishes target-speaker information from interference or uncertain components.

This reliability signal would directly control how incoming speech is routed: highly reliable target-speaker information could be integrated through channel fusion for stronger semantic grounding, uncertain but potentially useful information could remain accessible through cross-attention, while unreliable or non-target components could be suppressed.

In this way, filtering and routing become part of a single adaptive mechanism rather than two separate stages.

Expected Contribution

This work would formulate full-duplex listening as a selective information-routing problem rather than only target-speaker detection or fixed fusion. It would test whether reliability-aware dynamic routing can reduce the grounding–robustness trade-off between existing integration strategies and provide a basis for jointly modelling acoustic reliability, interaction relevance, and ongoing generation context.

Footnotes

  1. IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems. Arxiv: https://arxiv.org/abs/2606.06559

  2. How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue. Arxiv: https://arxiv.org/abs/2605.10199