Goal
Determine whether the semantic-grounding / robustness trade-off reported by Lu et al. can be reproduced in a smaller controlled setting.
This milestone was intended to establish a grounded CF/XA testbed before studying incoming speech during ongoing generation and, eventually, selective routing.
Outcome
Stopped after grounding gate failure. The reproduction pipeline worked, but neither route established the generalizable speech-semantic grounding needed for the intended comparison. M1 is concluded with an inconclusive routing result; the research hypothesis remains unresolved.
Experiments
- E001 — Reproduction of Lu et al.Complete
- E002 — M1b-B0 Bottleneck DiagnosisComplete
- E003 — M1b-B1 LM Ability ValidationComplete
- E004—M1b-B2 Direct Speech GroundingComplete
- E005—M1b-B2 Follow-up E0 Tiny Causal Grounding DiagnosisComplete
- E005—M1b-B2 Follow-up E1 Frozen-Language Interface DiagnosisComplete
The linked E001–E006 notes are the historical experiment record and the source of truth for individual results. Their next-experiment sections preserve plans made at the time; the final milestone decision is recorded here.
Findings
Initial reproduction and bottleneck diagnosis
E001 established an end-to-end CF/XA pipeline, but clean QA performance was weak and answers changed too little when speech was mismatched or removed. Its initial overlap evaluation could not distinguish robustness from ignoring incoming speech, and the answer-prefix setup further limited its discriminative power.
E002 found that the problem already existed before QA: native FastConformer ASR achieved approximately 12.57% WER, while the CF and XA ASR-stage interfaces into Qwen reached approximately 145.84% and 126.31% WER. Reliable speech-to-language alignment had never been established. Low QA loss could be driven by easy evidence continuation without correct question-dependent answer selection.
Encoder and language-model capabilities were independently validated
E003 separated the task and component capabilities. Context-only F1 was approximately 0.1075, correct-text QA reached 0.7601, and FastConformer ASR transcripts passed into the same QA model retained 0.7338 F1. The QA task was learnable, Qwen could solve it, and the encoder preserved useful linguistic information.
The main unresolved bottleneck was the direct path:
speech → FastConformer features → projector / CF / XA → QwenGrounding attempts did not generalize sufficiently
E004 increased ASR/QA exposure, but both routes remained near context-only performance and achieved 0/96 paired both-correct answers on tune. Speech swaps, zeroing, silence, and bypass controls showed little semantic dependence.
E005’s tiny-set diagnosis demonstrated genuine local speech-content learning for CF, but it missed most pre-registered fit targets and generalized poorly to the separate train-derived diagnostic examples. XA’s apparent fitting did not survive controls that removed speech content. This was evidence of partial local learnability for CF, not a successful grounding gate.
E006 froze the already capable B1 language model and trained the speech-side interface. CF showed limited but statistically supported speech-content dependence on tune, yet semantic accuracy and correct answer switching remained inadequate. XA remained largely insensitive to speech content. Both routes failed the complete grounding gate, with clean F1 around 0.10–0.15 and paired both-correct accuracy still zero.
These findings localize the problem mainly to speech-to-Qwen alignment under the tested setup. They do not establish that CF/XA is ineffective or that CF is scientifically superior to XA.
Conclusion / Decision
Conclude M1 and stop this experimental approach at the prerequisite capability gate. The attempted small-scale reproduction did not establish the grounding required for a meaningful CF-vs-XA comparison.
The later B3 interference stage and B4 were not pursued. If XA appeared to preserve generation better under interruption, that could simply reflect its failure to use the incoming speech. Without reliable grounding in both routes, such a result would not identify a robustness/grounding trade-off.
Increasing backbone size, widening XA, or substantially extending alignment training would have shifted the project into a larger speech-language alignment effort. Under the available compute and scope, this was a poor entry point for the intended routing question. The stopping gate prevented over-interpreting quantitative results from models that lacked the prerequisite capability.
Next
The research direction moves to Selective Perception Steering for Full-Duplex Spoken Dialogue, starting from an already pretrained full-duplex model, primarily PersonaPlex / Moshi. The plan is to freeze the underlying model initially, reproduce perception steering, test whether selective steering is worthwhile, assess whether hidden states predict when/how to intervene, and eventually train a lightweight controller.
This retains the question of when incoming speech should influence ongoing generation while beginning with existing speech competence. See the project-level decision and next direction for the planned E0–E3 progression and the lessons learned.
E006’s proposed direct-projector CTC follow-up remains preserved in that experiment note as a historical proposal; no result from that proposed experiment is claimed here.