CompleteDate: Sep 9, 2026

Question

Can the phenomenon in the UserStreamRoutingStudy2026 that CF has better semantic grounding while XA is more robust be reproduced in the smaller experiment?

Hypothesis

Semantic grounding:

As the interference increases, this gap will decrease and may eventually reverse, with XA becoming more robust than CF.

Description

Datasets

SQuAD (context, question, answer)

Training:

  • 300 ASR steps: learn the mapping from speech representation to text content

  • 600 QA steps: answer the spoken question based on the text context

  • 25% of QA examples contain strong overlap interference

Evaluation:

  • Noisy questions: degrade the original spoken question with different SNR levels

  • Mid-generation interference: inject another spoken question while the model is generating the answer to the first question

The dataset is constructed in pairs: the same paragraph contains two different questions with different answers. This allows testing whether the model actually uses the spoken question rather than guessing an answer from the context.

Network Architecture

        speech waveform
             ↓
    Frozen streaming FastConformer
             ↓
        512-d speech features
             ↓
        trainable projector
          512 → 1024
             ↓
        speech embeddings
             ↓
      ┌───────────────┐
      │               │
     CF              XA
      │               │
main stream      external memory
      │               │
      └── Qwen3-0.6B ──
                ↓
        Answer + Evidence

Qwen-CF and Qwen-XA are two models trained separately with the same backbone, training data and training budget.

Two seeds are used for each routing method:

CF seed17
CF seed42
XA seed17
XA seed42

Validity

Before comparing CF and XA scientifically, each model must first demonstrate that it actually uses the speech input.

The original validity gate was:

clean F1 ≥ 0.15
clean F1 - wrong-speech F1 ≥ 0.02

wrong-speech replaces the correct spoken question with another question from the same paragraph while keeping the expected answer unchanged.

The intuition is:

correct speech → correct question → correct answer
 
wrong speech → different question → performance should decrease

However, later auditing showed that this gate was too weak to establish reliable speech grounding.

Observed speech dependence:

ModelClean − Wrong F1Same answer for clean/wrong speech
CF seed170.020847/48
CF seed420.020845/48
XA seed170.011947/48
XA seed420.000047/48

For paired questions from the same paragraph:

ModelSame answer for two different questionsBoth questions answered correctly
CF seed1723/240/24
CF seed4220/240/24
XA seed1722/240/24
XA seed4221/240/24

Therefore, even though CF technically passed the original engineering gate, the models usually produced the same answer regardless of which spoken question was given.

This suggests that they primarily learned:

context → plausible answer

rather than:

spoken question
      ↓
understand what is being asked
      ↓
select the corresponding information from context
      ↓
answer

CF’s clean-vs-wrong gain also had a bootstrap 95% CI including zero, so the effect is not sufficiently reliable to claim successful speech grounding.

Scientific Test

Noise

The original spoken question is degraded using additive noise:

clean
20 dB SNR
10 dB SNR
0 dB SNR

The goal is to measure how CF and XA respond when the reliability of the incoming speech decreases.

Expected hypothesis:

CF performs better under clean speech,
but its advantage decreases as speech becomes less reliable.

Overlap

The model first receives Question A and begins answering it.

After 8 gold target tokens, Question B is injected as interfering user speech.

Interference durations:

1.28 s
2.56 s
5.12 s

The intended goal is to test whether XA preserves the current generation context better than CF when new conflicting user speech arrives.

However, the M1a interference task was later found to have weak discriminative power:

  • Across all four models and 192 paired cases, clean and strongest-overlap generations were identical.

  • 34/48 answers were already completely contained within the first 8 gold prefix tokens.

  • About 96.5% of evaluated continuation tokens could already access the new speech memory, so the lack of change cannot simply be explained by interference arriving too late.

  • After receiving the correct answer prefix, the model often only needed to continue copying the evidence sentence.

Therefore, identical outputs under interference cannot distinguish between:

Model understands the new speech
→ judges it irrelevant
→ deliberately ignores it

and:

Model does not meaningfully use the user speech stream
→ output remains unchanged

Result

M1a successfully established an end-to-end experimental pipeline for both CF and XA:

speech
→ streaming encoder
→ projector
→ CF / XA routing
→ Qwen
→ training
→ cached generation
→ evaluation

Both architectures can be trained, gradients flow correctly, tiny-set overfitting succeeds, and the full evaluation pipeline runs.

However, the main scientific hypothesis cannot yet be tested reliably.

The strongest problem is weak speech grounding.

CF seed17, for example, obtained:

clean F1 = 0.1636
wrong F1 = 0.1428
zero F1  = 0.1460

Although this technically passed the original validity threshold, the detailed audit showed that almost all clean/wrong inputs generated exactly the same answer.

XA showed the same overall problem.

Therefore:

M1a result ≠ CF/XA understand speech but perform poorly
 
M1a result = CF/XA can be trained,
             but reliable question-dependent speech grounding
             has not yet been established

The overlap experiment also cannot currently support a robustness claim because unchanged generations may simply result from the model ignoring the user stream.

No reliable CF-vs-XA trade-off can therefore be concluded from M1a.

Interpretation

The experiment did not yet reproduce the phenomenon reported by Lu et al.

The primary limitation is not simply low overall F1, but that the measured performance cannot be confidently attributed to understanding the spoken question.

The models appear to exploit a strong shortcut:

read context
→ choose a plausible answer/evidence span

while the speech input only weakly affects the semantic choice.

Several possible causes remain:

  • insufficient speech-to-LLM alignment training;

  • ASR alignment being forgotten during QA training;

  • the Answer + Evidence objective allowing context-copying to dominate the loss;

  • insufficient training data/exposure;

  • XA’s external-memory pathway being too weak or poorly optimized;

  • limited backbone capacity.

M1a cannot yet determine which of these is the dominant bottleneck.

Therefore, immediately switching from Qwen3-0.6B to Qwen3-1.7B or simply increasing XA capacity would be premature. A larger model could improve the score without resolving whether the system actually uses speech.

The main lesson from M1a is:

Before comparing CF and XA routing behavior, the experiment must first establish that both models reliably use the spoken question to select the correct answer.

Similarly, before interpreting resistance to interference as robustness, the experiment must demonstrate that the model is capable of using incoming speech when that speech is relevant.

Next Experiment

Goal

Establish reliable speech grounding for both CF and XA and redesign the interference task so that robustness can be distinguished from simply ignoring the user stream.

The next stage should first identify whether the current bottleneck lies in the acoustic frontend, speech-to-LLM alignment, training objective, XA pathway, or backbone capacity.

Expected Outcomes

A successful next stage should produce:

  1. CF and XA models whose answers clearly change according to which spoken question is provided.

  2. A substantial and statistically reliable gap between correct-speech and wrong/zero-speech conditions.

  3. An interference task where irrelevant speech can be ignored but relevant incoming speech can measurably influence generation.

  4. Enough diagnostic evidence to determine whether increasing XA capacity or upgrading the backbone to Qwen3-1.7B is actually necessary.