TL;DR
- Audio itself is fine. FastConformer can recover the question.
- Speech-to-Qwen alignment is not fine. CF/XA cannot reliably transcribe even at the ASR checkpoint.
- QA loss was misleading. Easy evidence copying hides failure in answer selection.
- XA is weak, but not yet proven under-capacity. The common alignment failure must be fixed first.
Question
Why did M1a fail to establish reliable speech grounding?
Specifically, where is information being lost in the pipeline?
speech
↓
FastConformer
↓
speech features
↓
projector + CF / XA
↓
Qwen
↓
answerPossible bottlenecks include:
- acoustic frontend quality;
- speech-to-LLM alignment;
- forgetting during QA training;
- QA objective / semantic selection;
- weak XA memory pathway;
- caching, timing, or numerical implementation issues.
Hypothesis
If the frozen speech encoder already preserves linguistic information, but CF/XA cannot transcribe the same speech, then the main bottleneck lies in speech-to-LLM alignment rather than the acoustic frontend.
If ASR ability is present after the ASR stage but disappears after QA training, catastrophic forgetting is likely.
If XA’s memory pathway has negligible numerical effect, XA may additionally suffer from weak routing capacity or optimization.
Description
B0 does not retrain the models. It audits the existing M1a checkpoints and pipeline.
Acoustic Frontend Check
Use the frozen FastConformer’s original ASR head to transcribe the spoken questions:
speech
↓
FastConformer
↓
original ASR head
↓
transcriptThis tests whether the encoder features still contain recoverable linguistic information.
Additional checks verify:
- cached feature recomputation;
- causal streaming behavior;
- absence of future-speech leakage.
Speech-to-LLM Alignment Check
Evaluate CF/XA as ASR systems without QA context:
speech
↓
FastConformer features
↓
projector
↓
CF / XA
↓
Qwen
↓
transcriptCompare the ASR-stage checkpoint with the final QA checkpoint.
This distinguishes:
alignment never learnedfrom:
alignment learned
↓
forgotten during QASemantic Selection Check
Separate the QA objective into:
Answer loss
Evidence lossand inspect whether changing the spoken question changes the selected answer.
This tests whether the low total QA loss observed in M1a actually reflected question understanding.
XA Pathway Check
Audit whether the external speech memory meaningfully affects the backbone.
Checks include:
- XA residual magnitude relative to the backbone hidden state;
- parameter updates and gradients;
- memory swap / bypass behavior;
- numerical behavior under BF16 versus FP32.
Result
Acoustic Frontend
The frozen FastConformer achieved:
Encoder WER = 12.57%This indicates that the speech itself is reasonably intelligible and that the encoder preserves linguistic information.
Cached recomputation and future-speech causality checks also passed.
Therefore, corrupted audio, bad feature extraction, and incorrect streaming timing are unlikely to be the main causes of M1a failure.
Speech-to-LLM Alignment
At the end of the ASR training stage:
CF ASR WER = 145.84%
XA ASR WER = 126.31%In comparison:
FastConformer native ASR WER = 12.57%The same speech therefore contains recoverable linguistic information, but CF/XA cannot reliably convert the encoder representation into language understood by Qwen.
The main failure occurs around:
FastConformer features
↓
projector + routing
↓
Qwen language representation
✗Importantly, this failure already exists at the ASR-stage checkpoint.
Therefore, QA-stage forgetting cannot be the main explanation.
The speech-to-LLM alignment was never sufficiently established in the first place.
Semantic Selection
All four final models achieved:
0 / 24paired paragraphs where both different spoken questions were answered correctly.
The loss audit also showed:
Evidence loss << Answer lossThe evidence sequence is relatively easy to predict or copy from the context, while selecting the correct answer remains difficult.
Therefore, the low total QA training loss observed in M1a was misleading.
The model could reduce its loss through:
context
↓
easy evidence continuationwithout learning:
spoken question
↓
semantic selection
↓
correct answerThis suggests that Answer + Evidence is not a clean training objective for measuring grounding.
XA Pathway
The XA pathway is not completely inactive.
Speech memory produces measurable residual changes and the pathway receives parameter updates, but its effect on the backbone is weak.
Therefore, the current evidence does not support:
XA is completely disconnectedHowever, it also does not yet support:
XA capacity is too smallbecause CF suffers from the same fundamental speech-alignment failure.
A numerical issue was also observed:
strict BF16 equivalence check → failed
FP32 re-check → discrepancy became much smallerThis suggests that some observed differences were caused by numerical precision rather than architectural errors.
Future diagnostics should therefore use a fixed precision and reference computation scheme.
Implementation Validation
12 CPU diagnostic tests passed.
All 192 clean generations matched the previous M1a outputs.
No model weights were changed and the held-out final test set was not evaluated.
Interpretation
B0 substantially narrows the failure location found in M1a.
The current pipeline can be summarized as:
Audio
↓
FastConformer
↓
✓ speech information preserved
↓
projector + CF / XA + Qwen
↓
✗ speech-language alignment not established
↓
QA answer selection
↓
✗ unreliableThe strongest current evidence points to speech-to-LLM alignment as the primary bottleneck.
The results make two previous explanations less likely:
- The speech encoder cannot recover the question.
- The model originally learned ASR but forgot it during QA training.
Instead, the ASR-stage models themselves already fail to reliably decode the speech representation.
The M1a QA objective also obscured this problem because long evidence generation could dominate the training loss even when answer selection remained poor.
For XA, the external-memory pathway exists but remains weak. Whether this is caused by capacity, optimization, initialization, or numerical precision is still unclear.
Therefore, B0 does not yet justify either:
Qwen3-0.6B → Qwen3-1.7Bor:
increase XA capacityas the first intervention.
The immediate priority is to construct a cleaner task where successful training requires using the question content and to establish a matched text baseline before changing model scale.
Next Experiment
Goal
Build a cleaner grounding task that directly tests question-dependent answer selection and determine whether Qwen3-0.6B can learn the same task when the question is provided as text.
Expected Outcomes
B1 should establish:
- a high-quality paired-question dataset with an answer-only objective;
- a matched text-QA baseline for the current backbone;
- clear separation between context-only and question-conditioned performance;
- enough evidence to determine whether the next bottleneck is alignment, XA capacity, or backbone scale.
Context-only low → question is necessray, no shortcut
Text Question + Context high → 0.6B Qwen can learn QA
ASR transcript + Context high-ish → speech frontend is good enough