E1 tested whether the speech interface can learn grounding when the already capable B1 text-QA model is frozen, removing simultaneous language-model adaptation as a possible source of failure.
CF showed limited but statistically supported speech-content dependence on tune: clean speech outperformed same-length wrong speech by +0.0785 F1, with 95% CI [0.0320, 0.1284].
However, CF still failed semantic grounding: clean F1 was only 0.155, Pair EM was 0, candidate switching was 4.17%, and free correct switching was 0.
XA showed almost no semantic dependence on speech content. Replacing clean speech with same-length wrong speech left 188/192 outputs unchanged, and silence left 183/192 unchanged.
Both routes also showed poor free ASR after the ASR stage, indicating that the speech-to-language mapping itself remains weak.
Since the frozen B1 text model reaches approximately 0.76 tune F1, while speech-conditioned CF/XA remain near 0.10–0.15, the bottleneck is increasingly localized to:
FastConformer representation↓projector + routing interface↓representation usable by Qwen
E1 therefore rules out the hypothesis that simply freezing a capable language model is sufficient. The remaining problem is more specifically speech-language alignment / interface learning.
The next preregistered experiment is E2, which adds direct-projector CTC supervision to explicitly align projected speech representations with Qwen’s lexical space.
B3 remains blocked.
Previous Experiment
E005 Tiny Overfit Testing asked whether CF and XA could at least learn speech semantics on a tiny fitted dataset.
E0 found:
CF→ strong fitted QA performance→ clear speech-content dependence→ poor generalizationXA→ apparent QA fitting→ almost no genuine speech-content dependence
For CF:
Fit F1 = 0.844Pair EM = 0.688Candidate switching = 0.750Clean − zero F1 = +0.310
This showed that CF was not fundamentally disconnected from speech.
However, performance dropped sharply on diagnostic questions:
Diagnostic F1 = 0.219Pair EM = 0Candidate switching = 0
XA showed even weaker evidence of semantic grounding.
The key remaining ambiguity was:
Maybe the interface is learnable,but jointly adapting the speech interfaceand language model makes optimization unstable.
E1 therefore fixes the language side and asks the speech interface to adapt to a stable target representation.
Question
Can CF and XA learn reliable speech grounding when the language model is already capable of the QA task and remains frozen?
More specifically:
Can a fresh speech interface learn to communicate with a fixed Qwen representation space?
Does correct speech produce better answers than wrong or removed speech?
Does changing the spoken question cause the answer to switch correctly?
Was simultaneous language-model adaptation a major cause of the previous failures?
Can both routes establish sufficient grounding to proceed toward B3?
Hypothesis
B1 already established that the language model can solve the QA task from text:
If the main problem in previous experiments was simultaneous language/interface adaptation, then fixing the target language space should make speech grounding easier.
Do not only tell the speech interface what final answer Qwen should produce. Explicitly force the projected speech representation to contain lexical information that Qwen’s own vocabulary space can decode.
E1’s CE-only QA branch remains the matched control.
Possible outcomes:
CTC improves+grounding improves↓weak speech-language alignment was a major bottleneck
or:
CTC improves+grounding still fails↓lexical information exists,but routing / Qwen readout cannot use it
or:
CTC itself fails↓the speech-to-language interface problemis even more fundamental
Immediate Goal
Do not enter B3.
The immediate objective remains:
Establish reliable and causal clean-speech grounding for both CF and XA before studying how routing should behave under interference.