Concluded — prerequisite speech grounding was not established; the selective routing hypothesis remains unresolved.

ConcludedStarted: Sep 2026

Timeline

Scroll or drag to explore milestones. Select a milestone to open it.

Overview / Research Question

Investigating how incoming user speech should influence an AI while it is already speaking.

Can a full-duplex spoken dialogue model selectively route incoming user speech according to its reliability and interaction relevance?

The project was motivated by the trade-off reported by Lu et al.: Channel Fusion (CF) integrates incoming speech more directly into the autoregressive stream, potentially strengthening semantic grounding while increasing interference; Cross-Attention (XA) keeps speech more separate and accesses it through cross-attention, potentially preserving ongoing generation better while providing weaker grounding.

The intended direction was dynamic, selective routing beyond fixed CF or XA. M1 — Reproduce CF vs. XA Behaviour first attempted a smaller controlled reproduction of that trade-off using a frozen FastConformer encoder, Qwen3-0.6B, and learned speech interfaces.

Outcome

Concluded — prerequisite speech grounding was not established.

The attempted small-scale reproduction did not establish the speech-language grounding required for a meaningful CF-vs-XA comparison. The original routing hypothesis remains unresolved. These results are neither evidence that CF/XA is ineffective nor evidence against selective user-stream routing itself.

Under the available compute and project scope, starting from a text LLM, a frozen speech encoder, and a lightweight learned interface proved to be a poor experimental entry point for this question.

Why the Reproduction Stopped

The initial pipeline ran end-to-end, but semantic behaviour was suspiciously weak: correct, mismatched, or absent audio often produced nearly the same answers. Diagnosis then showed that the speech encoder retained useful linguistic information and Qwen could solve the QA task from text, while the direct speech-feature-to-Qwen alignment remained unreliable. Longer training, tiny-set fitting, and freezing an already capable language model did not establish generalizable grounding.

The pre-registered grounding gate was never passed. The later B3 interference stage would have injected incoming user speech during ongoing generation, but without this prerequisite its results would have been uninterpretable. Apparent XA robustness could simply mean that XA ignored speech. E001’s initial overlap test had already exposed this ambiguity; it did not establish a robustness advantage.

The project therefore stopped before B3/B4. Scaling Qwen, widening XA, or committing to substantially more alignment training would have expanded the work into a separate speech-language alignment engineering effort. That capability was a prerequisite for the research question, and rebuilding it would have displaced the question itself.

Key Findings

  • E001 — Working pipeline, unreliable grounding. Both routes trained and ran through evaluation, but weak question-dependent behaviour prevented a meaningful reproduction of the CF/XA trade-off.
  • E002 — The bottleneck preceded QA. Native encoder ASR worked, while CF/XA transcription through Qwen collapsed even at the ASR checkpoint. Alignment had not been reliably learned; easy evidence copying could reduce QA loss without spoken-question understanding.
  • E003 — The components worked separately. Correct-text QA and encoder-ASR-transcript-to-QA performed strongly compared with context-only QA. This validated the task, Qwen’s QA ability, and the encoder’s usable linguistic information, localizing the main bottleneck to speech features → projector / CF / XA → Qwen.
  • E004–E006 — More fitting did not establish generalizable grounding. Longer training and a frozen, capable text-QA model did not pass the gate. CF showed local speech-content learning and later limited but statistically supported content dependence on tune examples; this still fell short of reliable semantic accuracy. XA remained largely insensitive to speech content. Neither route supported the intended interference comparison.

The linked experiment notes retain the detailed metrics, controls, limitations, and decisions made at each stage.

Lessons Learned

  1. Validate prerequisites before the high-level hypothesis. Before comparing CF and XA under interruption, demonstrate that both routes understand speech.
  2. Architectural similarity is not enough for reproduction. A smaller frozen backbone with lightweight adapters may retain the names CF and XA without preserving the speech competence needed for the original phenomenon.
  3. Training loss is not grounding evidence. Correct, mismatched, zero, and removed speech controls, paired questions, and held-out evaluation were more informative than loss reduction.
  4. Factorized ablations localize bottlenecks. Separating encoder ability, text-QA ability, ASR→text-QA ability, and direct speech→LLM ability avoided attributing every failure to the backbone or model size.
  5. A stopping criterion protects interpretation. The grounding gate prevented proceeding to B3/B4 and drawing routing conclusions from models that lacked the necessary capability.
  6. Keep prerequisite engineering within scope. Do not spend most of a research project rebuilding a capability needed only to begin testing the actual question. This lesson motivated the next direction.

Decision / Next Direction

The investigation is concluded at the grounding gate. Historical experiment proposals remain in E001–E006 as records of the decisions at the time; they are not an active continuation plan. In particular, E006’s proposed direct-projector CTC alignment is a proposal, not an additional reported result.

The next project is Selective Perception Steering for Full-Duplex Spoken Dialogue. The broader question remains: how should incoming user speech influence ongoing full-duplex generation, and when should the model prioritize listening over continuing to speak?

The experimental entry point changes to an already pretrained full-duplex spoken dialogue model, primarily PersonaPlex / Moshi. The plan is to keep the model frozen initially and intervene directly in its internal representations, studying a direction from a more generative/speaking state toward a more perceptive/listening state:

The question is whether steering strength should vary dynamically with incoming speech.

InvestigationExperimental progression
Selective User-Stream RoutingSpeech → encoder → learn speech-to-LLM alignment through CF/XA → establish grounding → study selective routing
Selective Perception SteeringPretrained full-duplex model with existing speech competence → reproduce perception steering → test when it helps → predict when to steer → learn a lightweight controller

The planned stages are:

  • E0 — Steering can work. Demonstrate perception steering on a frozen full-duplex model, especially under interruption.
  • E1 — Selective steering is worth pursuing. Use oracle conditions and steering strengths to test whether selective or dynamic steering improves on always-on steering.
  • E2 — The model can know when/how to steer. Test whether hidden states contain enough information to predict relevance, reliability, or the optimal intervention.
  • E3 — Automatic selective steering improves behaviour. Train a lightweight controller to choose steering strength dynamically while keeping the underlying model frozen.

These are planned stages for the new project, distinct from the E0/E1 follow-up labels within this project’s B2 notes. The aim is to preserve the selective-listening question while starting from existing speech competence.

Experiments

Open M1 — Reproduce CF vs. XA Behaviour for the milestone synthesis. The individual notes remain the source of truth for each experiment: