CompleteDate: Sep 12, 2026

TL;DR

  1. The new answer-only paired-question task requires question information: context-only reaches only 0.1075 F1.

  2. Qwen3-0.6B can learn the task when given the correct text question: 0.7601 F1.

  3. Frozen FastConformer → ASR transcript → the same QA model retains most performance: 0.7338 F1.

  4. Therefore, the acoustic frontend and 0.6B backbone are not the main bottlenecks. The remaining problem is primarily the direct speech-to-LLM alignment through CF/XA.

Previous Experiment

E002—M1b-B0 Bottleneck Diagnosis showed that M1a failed mainly because speech-to-Qwen alignment was never reliably established.

The frozen FastConformer could recover the spoken question with reasonable ASR quality, while CF/XA could not reliably transcribe the same encoder representation even at the ASR-stage checkpoint.

It also showed that the old Answer + Evidence objective was misleading because evidence copying could reduce training loss without requiring correct question-dependent answer selection.

However, B0 still left two important possibilities unresolved:

Maybe Qwen3-0.6B itself cannot learn this QA task
 
or
 
Maybe useful question information is lost before it reaches the language model

B1 is designed to isolate these possibilities before modifying CF/XA capacity or backbone size.

Question

Can the current testbed cleanly measure question-dependent grounding?

More specifically:

  1. Does the task actually require knowing the question?

  2. Can Qwen3-0.6B solve the task when the question is provided correctly?

  3. Does the frozen speech encoder preserve enough linguistic information for the same QA task?

Hypothesis

If:

Context only
→ poor performance

but:

Context + text question
→ strong performance

then the task genuinely requires question information and Qwen3-0.6B is capable of solving it.

If:

speech
↓
FastConformer
↓
ASR transcript
↓
same text QA model

also performs well, then useful question information survives the acoustic frontend.

In that case, the remaining bottleneck should lie mainly in:

FastConformer representation
↓
CF / XA
↓
Qwen

rather than in the speech encoder or backbone capacity.

Description

B1 does not train CF or XA.

It compares three controlled baselines on the same answer-only QA task.

A. Context-only

Context
↓
Qwen
↓
Answer

The question is empty.

This measures how much can be solved using context shortcuts alone.

B. Correct Text Question

Context + correct text question
↓
Qwen
↓
Answer

This measures whether Qwen3-0.6B can learn the QA task under the same 600-step QA budget planned for B2 R0.

C. Encoder-ASR Cascade

Spoken question
↓
Frozen FastConformer
↓
CTC transcript
↓
same QA checkpoint as B
↓
Answer

No additional QA model is trained.

This tests whether the frozen encoder preserves enough linguistic information for downstream question answering.

The evaluation contains:

192 questions
96 context pairs

with two different questions for each context.

A pair is only counted as fully correct when both questions are answered correctly.

Result

Context-only Baseline

F1 = 0.1075
Pair EM = 0 / 96

The model produced the same answer for both questions in all 96 context pairs.

Therefore, context alone is not sufficient to solve the task.

The new testbed successfully requires question-dependent information.

Text Question Baseline

F1 = 0.7601
Pair EM = 30 / 96

Compared with context-only:

ΔF1 = +0.6526
95% CI = [0.5904, 0.7108]

The large improvement shows that Qwen3-0.6B can learn the task when the question is available.

Therefore:

Qwen3-0.6B is too weak

is no longer a good explanation for the poor M1a CF/XA results.

Encoder-ASR Cascade

F1 = 0.7338
Pair EM = 27 / 96

The frozen encoder achieved:

WER = 11.53%
CER = 4.45%

Compared with the correct-text baseline:

Text QA     = 0.7601
ASR → QA    = 0.7338
 
ΔF1 = -0.0263

Most of the text-QA performance is therefore preserved after passing the spoken question through the frozen speech encoder and CTC decoder.

This strongly suggests that the encoder representation still contains the linguistic information required for question answering.

The remaining small performance loss cannot be interpreted purely as acoustic information loss because transcript formatting, punctuation, casing, tokenization, and ASR errors can all affect the downstream language model.

Interpretation

B1 substantially narrows the bottleneck identified in E002 Bottleneck Diagnosis.

The current evidence now looks like:

speech
↓
FastConformer
↓
✓ question information preserved
↓
CTC transcript
↓
Qwen
↓
✓ QA works

but the historical M1a path was:

speech
↓
FastConformer
↓
speech representation
↓
CF / XA
↓
Qwen
↓
✗ unreliable question grounding

This means that neither of the following is currently the strongest explanation:

speech encoder is too weak

or:

Qwen3-0.6B is too weak

Instead, the remaining problem is increasingly localized to the direct speech-to-language interface:

FastConformer features
        ↓
projector + CF / XA
        ↓
Qwen representation
        ?

B1 therefore does not justify immediately upgrading:

Qwen3-0.6B → Qwen3-1.7B

or increasing XA capacity.

The first priority is to determine whether CF and XA can establish basic speech grounding under the cleaner B1 task and a matched training recipe.

Next Experiment

B2 — CF/XA Grounding and Training Recipe

Goal

Test whether CF and XA can directly convert the preserved FastConformer speech representation into information that Qwen can use for question-dependent answer generation.

First validate the actual CF/XA FP32/reference paths.

Then run the common baseline:

R0
 
300 ASR steps
+
600 answer-only QA steps

for both CF and XA under matched:

  • data;

  • training order;

  • initialization;

  • optimization;

  • frozen encoder;

  • Qwen3-0.6B backbone.

The main question is no longer whether the task is learnable.

It is:

Can CF/XA bridge the gap between the speech representation and the language model?

Desired evidence:

same context
+
different spoken questions
↓
different correct answers

If R0 fails, later B2 recipes can test whether the cause is insufficient alignment training, ASR forgetting, or route-specific optimization.

Only after basic grounding is established should XA capacity or backbone scale be changed.