TL;DR
The new answer-only paired-question task requires question information: context-only reaches only 0.1075 F1.
Qwen3-0.6B can learn the task when given the correct text question: 0.7601 F1.
Frozen FastConformer → ASR transcript → the same QA model retains most performance: 0.7338 F1.
Therefore, the acoustic frontend and 0.6B backbone are not the main bottlenecks. The remaining problem is primarily the direct speech-to-LLM alignment through CF/XA.
Previous Experiment
E002—M1b-B0 Bottleneck Diagnosis showed that M1a failed mainly because speech-to-Qwen alignment was never reliably established.
The frozen FastConformer could recover the spoken question with reasonable ASR quality, while CF/XA could not reliably transcribe the same encoder representation even at the ASR-stage checkpoint.
It also showed that the old Answer + Evidence objective was misleading because evidence copying could reduce training loss without requiring correct question-dependent answer selection.
However, B0 still left two important possibilities unresolved:
Maybe Qwen3-0.6B itself cannot learn this QA task
or
Maybe useful question information is lost before it reaches the language modelB1 is designed to isolate these possibilities before modifying CF/XA capacity or backbone size.
Question
Can the current testbed cleanly measure question-dependent grounding?
More specifically:
-
Does the task actually require knowing the question?
-
Can Qwen3-0.6B solve the task when the question is provided correctly?
-
Does the frozen speech encoder preserve enough linguistic information for the same QA task?
Hypothesis
If:
Context only
→ poor performancebut:
Context + text question
→ strong performancethen the task genuinely requires question information and Qwen3-0.6B is capable of solving it.
If:
speech
↓
FastConformer
↓
ASR transcript
↓
same text QA modelalso performs well, then useful question information survives the acoustic frontend.
In that case, the remaining bottleneck should lie mainly in:
FastConformer representation
↓
CF / XA
↓
Qwenrather than in the speech encoder or backbone capacity.
Description
B1 does not train CF or XA.
It compares three controlled baselines on the same answer-only QA task.
A. Context-only
Context
↓
Qwen
↓
AnswerThe question is empty.
This measures how much can be solved using context shortcuts alone.
B. Correct Text Question
Context + correct text question
↓
Qwen
↓
AnswerThis measures whether Qwen3-0.6B can learn the QA task under the same 600-step QA budget planned for B2 R0.
C. Encoder-ASR Cascade
Spoken question
↓
Frozen FastConformer
↓
CTC transcript
↓
same QA checkpoint as B
↓
AnswerNo additional QA model is trained.
This tests whether the frozen encoder preserves enough linguistic information for downstream question answering.
The evaluation contains:
192 questions
96 context pairswith two different questions for each context.
A pair is only counted as fully correct when both questions are answered correctly.
Result
Context-only Baseline
F1 = 0.1075
Pair EM = 0 / 96The model produced the same answer for both questions in all 96 context pairs.
Therefore, context alone is not sufficient to solve the task.
The new testbed successfully requires question-dependent information.
Text Question Baseline
F1 = 0.7601
Pair EM = 30 / 96Compared with context-only:
ΔF1 = +0.6526
95% CI = [0.5904, 0.7108]The large improvement shows that Qwen3-0.6B can learn the task when the question is available.
Therefore:
Qwen3-0.6B is too weakis no longer a good explanation for the poor M1a CF/XA results.
Encoder-ASR Cascade
F1 = 0.7338
Pair EM = 27 / 96The frozen encoder achieved:
WER = 11.53%
CER = 4.45%Compared with the correct-text baseline:
Text QA = 0.7601
ASR → QA = 0.7338
ΔF1 = -0.0263Most of the text-QA performance is therefore preserved after passing the spoken question through the frozen speech encoder and CTC decoder.
This strongly suggests that the encoder representation still contains the linguistic information required for question answering.
The remaining small performance loss cannot be interpreted purely as acoustic information loss because transcript formatting, punctuation, casing, tokenization, and ASR errors can all affect the downstream language model.
Interpretation
B1 substantially narrows the bottleneck identified in E002 Bottleneck Diagnosis.
The current evidence now looks like:
speech
↓
FastConformer
↓
✓ question information preserved
↓
CTC transcript
↓
Qwen
↓
✓ QA worksbut the historical M1a path was:
speech
↓
FastConformer
↓
speech representation
↓
CF / XA
↓
Qwen
↓
✗ unreliable question groundingThis means that neither of the following is currently the strongest explanation:
speech encoder is too weakor:
Qwen3-0.6B is too weakInstead, the remaining problem is increasingly localized to the direct speech-to-language interface:
FastConformer features
↓
projector + CF / XA
↓
Qwen representation
?B1 therefore does not justify immediately upgrading:
Qwen3-0.6B → Qwen3-1.7Bor increasing XA capacity.
The first priority is to determine whether CF and XA can establish basic speech grounding under the cleaner B1 task and a matched training recipe.
Next Experiment
B2 — CF/XA Grounding and Training Recipe
Goal
Test whether CF and XA can directly convert the preserved FastConformer speech representation into information that Qwen can use for question-dependent answer generation.
First validate the actual CF/XA FP32/reference paths.
Then run the common baseline:
R0
300 ASR steps
+
600 answer-only QA stepsfor both CF and XA under matched:
-
data;
-
training order;
-
initialization;
-
optimization;
-
frozen encoder;
-
Qwen3-0.6B backbone.
The main question is no longer whether the task is learnable.
It is:
Can CF/XA bridge the gap between the speech representation and the language model?
Desired evidence:
same context
+
different spoken questions
↓
different correct answersIf R0 fails, later B2 recipes can test whether the cause is insufficient alignment training, ASR forgetting, or route-specific optimization.
Only after basic grounding is established should XA capacity or backbone scale be changed.