CompleteDate: Sep 11, 2026

TL;DR

  1. Audio itself is fine. FastConformer can recover the question.
  2. Speech-to-Qwen alignment is not fine. CF/XA cannot reliably transcribe even at the ASR checkpoint.
  3. QA loss was misleading. Easy evidence copying hides failure in answer selection.
  4. XA is weak, but not yet proven under-capacity. The common alignment failure must be fixed first.

Question

Why did M1a fail to establish reliable speech grounding?

Specifically, where is information being lost in the pipeline?

speech
↓
FastConformer
↓
speech features
↓
projector + CF / XA
↓
Qwen
↓
answer

Possible bottlenecks include:

  • acoustic frontend quality;
  • speech-to-LLM alignment;
  • forgetting during QA training;
  • QA objective / semantic selection;
  • weak XA memory pathway;
  • caching, timing, or numerical implementation issues.

Hypothesis

If the frozen speech encoder already preserves linguistic information, but CF/XA cannot transcribe the same speech, then the main bottleneck lies in speech-to-LLM alignment rather than the acoustic frontend.

If ASR ability is present after the ASR stage but disappears after QA training, catastrophic forgetting is likely.

If XA’s memory pathway has negligible numerical effect, XA may additionally suffer from weak routing capacity or optimization.

Description

B0 does not retrain the models. It audits the existing M1a checkpoints and pipeline.

Acoustic Frontend Check

Use the frozen FastConformer’s original ASR head to transcribe the spoken questions:

speech
↓
FastConformer
↓
original ASR head
↓
transcript

This tests whether the encoder features still contain recoverable linguistic information.

Additional checks verify:

  • cached feature recomputation;
  • causal streaming behavior;
  • absence of future-speech leakage.

Speech-to-LLM Alignment Check

Evaluate CF/XA as ASR systems without QA context:

speech
↓
FastConformer features
↓
projector
↓
CF / XA
↓
Qwen
↓
transcript

Compare the ASR-stage checkpoint with the final QA checkpoint.

This distinguishes:

alignment never learned

from:

alignment learned
↓
forgotten during QA

Semantic Selection Check

Separate the QA objective into:

Answer loss
Evidence loss

and inspect whether changing the spoken question changes the selected answer.

This tests whether the low total QA loss observed in M1a actually reflected question understanding.

XA Pathway Check

Audit whether the external speech memory meaningfully affects the backbone.

Checks include:

  • XA residual magnitude relative to the backbone hidden state;
  • parameter updates and gradients;
  • memory swap / bypass behavior;
  • numerical behavior under BF16 versus FP32.

Result

Acoustic Frontend

The frozen FastConformer achieved:

Encoder WER = 12.57%

This indicates that the speech itself is reasonably intelligible and that the encoder preserves linguistic information.

Cached recomputation and future-speech causality checks also passed.

Therefore, corrupted audio, bad feature extraction, and incorrect streaming timing are unlikely to be the main causes of M1a failure.

Speech-to-LLM Alignment

At the end of the ASR training stage:

CF ASR WER = 145.84%
XA ASR WER = 126.31%

In comparison:

FastConformer native ASR WER = 12.57%

The same speech therefore contains recoverable linguistic information, but CF/XA cannot reliably convert the encoder representation into language understood by Qwen.

The main failure occurs around:

FastConformer features
        ↓
projector + routing
        ↓
Qwen language representation
        ✗

Importantly, this failure already exists at the ASR-stage checkpoint.

Therefore, QA-stage forgetting cannot be the main explanation.

The speech-to-LLM alignment was never sufficiently established in the first place.

Semantic Selection

All four final models achieved:

0 / 24

paired paragraphs where both different spoken questions were answered correctly.

The loss audit also showed:

Evidence loss << Answer loss

The evidence sequence is relatively easy to predict or copy from the context, while selecting the correct answer remains difficult.

Therefore, the low total QA training loss observed in M1a was misleading.

The model could reduce its loss through:

context
↓
easy evidence continuation

without learning:

spoken question
↓
semantic selection
↓
correct answer

This suggests that Answer + Evidence is not a clean training objective for measuring grounding.

XA Pathway

The XA pathway is not completely inactive.

Speech memory produces measurable residual changes and the pathway receives parameter updates, but its effect on the backbone is weak.

Therefore, the current evidence does not support:

XA is completely disconnected

However, it also does not yet support:

XA capacity is too small

because CF suffers from the same fundamental speech-alignment failure.

A numerical issue was also observed:

strict BF16 equivalence check → failed
FP32 re-check → discrepancy became much smaller

This suggests that some observed differences were caused by numerical precision rather than architectural errors.

Future diagnostics should therefore use a fixed precision and reference computation scheme.

Implementation Validation

12 CPU diagnostic tests passed.

All 192 clean generations matched the previous M1a outputs.

No model weights were changed and the held-out final test set was not evaluated.

Interpretation

B0 substantially narrows the failure location found in M1a.

The current pipeline can be summarized as:

Audio
  ↓
FastConformer
  ↓
✓ speech information preserved
  ↓
projector + CF / XA + Qwen
  ↓
✗ speech-language alignment not established
  ↓
QA answer selection
  ↓
✗ unreliable

The strongest current evidence points to speech-to-LLM alignment as the primary bottleneck.

The results make two previous explanations less likely:

  1. The speech encoder cannot recover the question.
  2. The model originally learned ASR but forgot it during QA training.

Instead, the ASR-stage models themselves already fail to reliably decode the speech representation.

The M1a QA objective also obscured this problem because long evidence generation could dominate the training loss even when answer selection remained poor.

For XA, the external-memory pathway exists but remains weak. Whether this is caused by capacity, optimization, initialization, or numerical precision is still unclear.

Therefore, B0 does not yet justify either:

Qwen3-0.6B → Qwen3-1.7B

or:

increase XA capacity

as the first intervention.

The immediate priority is to construct a cleaner task where successful training requires using the question content and to establish a matched text baseline before changing model scale.

Next Experiment

Goal

Build a cleaner grounding task that directly tests question-dependent answer selection and determine whether Qwen3-0.6B can learn the same task when the question is provided as text.

Expected Outcomes

B1 should establish:

  1. a high-quality paired-question dataset with an answer-only objective;
  2. a matched text-QA baseline for the current backbone;
  3. clear separation between context-only and question-conditioned performance;
  4. enough evidence to determine whether the next bottleneck is alignment, XA capacity, or backbone scale.
Context-only               low              → question is necessray, no shortcut
Text Question + Context    high             → 0.6B Qwen can learn QA
ASR transcript + Context   high-ish         → speech frontend is good enough