TL;DR
CF and XA both failed to establish reliable direct speech grounding under the current B2 recipes.
Clean QA F1 remained around 0.11 for both routes, close to the B1 context-only baseline of 0.1075, while the matched text-question control reached 0.7730 F1.
Wrong, zero, silent, or bypassed speech produced almost the same answers as correct speech, and Pair EM was 0/96 for both routes.
Increasing ASR/QA training from R0 to R1 did not improve grounding, so insufficient training steps alone are unlikely to explain the failure.
The remaining bottleneck is now localized to the direct speech-to-Qwen interface.
FastConformer features
↓
projector
↓
CF / XA
↓
QwenB3 should not begin until basic speech-dependent QA is established.
Previous Experiment
E003 M1b-B1 Grounding Task Validation established that the new answer-only paired-question task is valid.
It showed:
Context only
→ F1 = 0.1075
Context + text question
→ F1 = 0.7601
Speech
↓
FastConformer
↓
CTC transcript
↓
same QA model
→ F1 = 0.7338Therefore:
-
Qwen3-0.6B can learn the QA task;
-
the task genuinely requires question information;
-
the frozen FastConformer preserves enough linguistic information for QA.
This narrowed the unresolved bottleneck to:
FastConformer speech representation
↓
projector + CF / XA
↓
QwenB2 directly tests whether this interface can establish speech grounding.
Question
Can CF and XA convert FastConformer speech representations into information that Qwen can use for question-dependent answer generation?
More specifically:
-
Can correct spoken questions improve QA performance above the context-only baseline?
-
Can the model distinguish two spoken questions that share the same context?
-
Does changing or removing the speech change the generated answer?
-
Is failure caused simply by insufficient ASR/QA training exposure?
-
Are CF and XA ready for the generation-interference experiments in B3?
Hypothesis
If direct speech grounding is learned, then:
Context
+
Speech Question A
↓
Answer A
Context
+
Speech Question B
↓
Answer BFor paired questions sharing the same context:
swap speech
↓
answer should also switchand removing the speech should cause a clear performance drop:
correct speech
>
wrong speech
>
zero / silence / bypassThe B2 development gate therefore requires both strong clean QA performance and clear causal dependence on speech content.
Description
B2 trains CF and XA directly on the B1 answer-only task.
Unlike B1:
text question
→ QwenB2 removes the question transcript from the prompt.
The only path through which Qwen can receive question information is:
Spoken question
↓
Frozen FastConformer
↓
speech features
↓
projector
↓
CF / XA
↓
Qwen
↓
AnswerThe text context is still given directly to Qwen.
Shared Setup
Both routes use:
-
Qwen3-0.6B;
-
frozen FastConformer;
-
the same 2,048 training questions;
-
192 tune questions / 96 context pairs;
-
answer-only QA supervision;
-
identical LoRA initialization;
-
identical projector initialization;
-
identical training order;
-
FP32 training;
-
no transcript input during QA inference.
CF and XA have approximately matched trainable parameter counts.
CF: 19.01M trainable parameters
XA: 19.30M trainable parametersR0 — Baseline Recipe
300 ASR updates
+
600 QA updatesBoth CF and XA are trained independently under the same exposure schedule.
R1 — Longer Alignment Training
Because R0 failed the development criteria, B2 activates R1:
1200 ASR updates
+
1200 QA updatesThe purpose is to test whether the R0 failure was mainly caused by insufficient training exposure.
R1 increases both speech-transcription training and QA training while keeping the route architecture unchanged.
Speech Dependency Controls
Final checkpoints are evaluated with:
clean
native wrong speech
same-length wrong speech
projected zero
silence
route bypassThe critical question is not only whether clean F1 is high.
The model must behave differently when the spoken question changes.
Result
R0
Clean QA
CF clean F1 = 0.1093
XA clean F1 = 0.1134
CF Pair EM = 0 / 96
XA Pair EM = 0 / 96For comparison:
B1 context-only = 0.1075
B1 text question = 0.7601Therefore, providing the spoken question through CF/XA produces almost no improvement over giving Qwen no question at all.
Speech Dependency
CF:
clean 0.1093
native wrong 0.1110
same-length wrong 0.1093
projected zero 0.1128
silence 0.1107
bypass 0.1128XA:
clean 0.1134
native wrong 0.1075
same-length wrong 0.1134
projected zero 0.1134
silence 0.1134
bypass 0.1134The model behaves almost identically regardless of whether the speech is:
correct
wrong
zero
silent
absentThe answer therefore shows almost no causal dependence on speech content.
Answer Switching
When the spoken question is replaced by the paired alternative question:
CF switching rate = 0%
XA switching rate = 0%A grounded model should change its answer when the question changes.
Neither route does so.
R0 therefore fails the B2 development gate.
R1
R1 substantially increases training exposure:
R0:
300 ASR + 600 QA
R1:
1200 ASR + 1200 QAHowever, final QA performance remains essentially unchanged.
CF clean F1 = 0.1081
XA clean F1 = 0.1111
CF Pair EM = 0 / 96
XA Pair EM = 0 / 96Speech dependency also remains weak.
CF:
clean 0.1081
native wrong 0.1193
same-length wrong 0.0953
projected zero 0.1015
silence 0.1035
bypass 0.1015XA:
clean 0.1111
native wrong 0.1074
same-length wrong 0.1111
projected zero 0.1111
silence 0.1111
bypass 0.1111Answer switching remains almost absent:
CF switching rate ≈ 1%
XA switching rate = 0%Therefore, simply increasing the R0 training budget does not establish speech grounding.
Matched Text Control
The matched R1 text-question control reaches:
F1 = 0.7730
Pair EM = 33 / 96The contrast is therefore approximately:
text question
→ F1 = 0.773
speech through CF/XA
→ F1 ≈ 0.11The underlying QA task remains easily learnable when the question reaches Qwen in textual form.
The failure is specific to the direct speech-conditioning path.
ASR Diagnostic
The route models also fail to reliably transcribe the question during free ASR decoding.
Final diagnostic WER remains around:
CF ≈ 0.92
XA ≈ 1.00despite longer ASR training.
At the same time, teacher-forced ASR loss decreases during training.
This suggests a possible shortcut:
correct transcript prefix
↓
Qwen learns how questions usually continuewithout necessarily learning:
speech features
↓
actual spoken contentTeacher-forced loss reduction therefore cannot be treated as evidence of speech grounding.
Free decoding and causal speech interventions remain necessary.
Why R2 Was Not Run
R2 would mix ASR and QA training to test whether QA training destroys speech recognition learned during the ASR stage.
However, R1 does not show the required ASR-forgetting pattern.
The final ASR WER is not substantially worse than the earlier ASR-stage diagnostic.
Therefore, the evidence does not support:
ASR grounding was learned
↓
QA training caused catastrophic forgettingInstead, the more likely problem is:
robust speech-to-Qwen grounding
was never establishedR2 was therefore not activated.
Why XA Capacity Was Not Increased
XA width128/256 is only meaningful if:
CF succeeds
but
XA width64 failsThat would indicate a route-specific XA capacity problem.
Instead:
CF fails
XA failsTherefore, the current evidence points to a shared speech-to-language alignment problem rather than an XA-specific bottleneck.
Increasing XA width would not resolve the scientific ambiguity.
Interpretation
B2 strengthens the bottleneck identified in B0 and B1.
The evidence now forms the following chain:
[E002]
Speech
↓
FastConformer
↓
✓ useful information exists
but
speech features
↓
CF / XA
↓
Qwen
↓
✗ unreliable groundingThen:
[E003]
text context + text question
↓
Qwen
↓
✓ QA works
Speech
↓
FastConformer
↓
CTC transcript
↓
Qwen
↓
✓ QA mostly worksNow B2 shows:
text context
+
Speech
↓
FastConformer
↓
projector
↓
CF / XA
↓
Qwen
↓
✗ QA remains near context-onlyThis makes the remaining bottleneck increasingly specific:
FastConformer features
↓
speech-language interface
↓
projector + CF / XA
↓
Qwen semantic computationThe current experiment does not show that CF or XA are fundamentally ineffective routing mechanisms.
It shows that:
Under the current training objective, initialization, and exposure budget, neither route learns to make Qwen depend reliably on spoken question semantics.
Therefore, B2 cannot be used to compare:
CF vs XAbecause neither model has passed the prerequisite of understanding the speech.
A route that ignores speech may appear robust to interference simply because the interference never affects generation.
Why B3 Is Blocked
B3 is intended to study incoming speech during ongoing generation.
For example:
Qwen generating answer
+
new user speech arrives
↓
CF / XA
↓
effect on ongoing generationThis comparison is only meaningful if both routes already understand speech.
Otherwise:
speech interference has little effectcould mean:
the route is robustor simply:
the route ignores speechB2 cannot distinguish these interpretations because both CF and XA currently fail basic grounding.
Therefore:
B3 = blockedNo B3/B4, confirmation/test inference, backbone upgrade, or M2 experiment should begin yet.
Conclusion
B2 rules out one simple explanation for the previous failure:
R0 was just too shortIncreasing the training budget substantially does not improve grounding.
Together with B1, the current evidence suggests:
Qwen QA ability ✓
FastConformer information ✓
Direct speech → Qwen ✗The next stage should therefore focus specifically on speech-language interface learnability and alignment, not on CF/XA scientific comparison.
Next Experiment
B2 Follow-up — Speech-to-Qwen Alignment Diagnosis
Step 1 — Tiny Causal Grounding Test
Before scaling anything, test whether the existing interface can overfit a very small dataset.
Use approximately:
8–16 context pairs
16–32 spoken questionswith both CF and XA.
Start from the validated B1 QA state and freeze Qwen.
Train only:
projector
+
CF / XA
+
necessary listening parametersusing the ordinary answer QA objective.
Success must require more than low training loss.
The model should satisfy:
Speech A → Answer A
Speech B → Answer B
swap A/B
→ answers also swap
zero / silence / bypass
→ performance collapsesand free ASR decoding should reflect actual speech content rather than only teacher-forced NLL improvement.
Decision
If the tiny set succeeds:
architecture is learnable
↓
investigate generalization / exposure / curriculumIf the tiny set fails:
basic interface learning is broken
↓
inspect representation and optimizationincluding:
-
projector output scale vs Qwen token embeddings;
-
projector initialization;
-
CF speech contribution magnitude;
-
XA gate and residual contribution;
-
route parameter update magnitude;
-
correct-vs-wrong speech logit differences;
-
speech-conditioned vs zero-conditioned hidden states.
Step 2 — Stronger Alignment Supervision
If Frozen-Qwen + ordinary QA CE still fails, test stronger objectives separately.
First candidate:
Text Teacher
context + text question
↓
validated B1 Qwen
↓
next-token distribution
│
│ KD
▼
Speech Student
context + speech
↓
FastConformer
↓
projector + CF / XA
↓
same frozen QwenTrain with:
QA loss
+
speech/text behavioral distillationThe transcript may be used by the teacher during training but must never enter the speech student’s inference input.
If this still fails, separately test a more direct auxiliary speech-semantic objective such as:
CTC-style transcript alignmentor:
speech-conditioned hidden representation
≈
text-conditioned hidden representationEach intervention should be tested independently so that the cause of any improvement remains identifiable.
Do Not Scale Yet
Until tiny causal grounding succeeds, do not:
Qwen3-0.6B → Qwen3-1.7B
XA width64 → width128/256
massively increase training steps
enter B3The immediate goal is simpler:
Make Qwen’s answer causally depend on what was spoken.