CompleteDate: Sep 13, 2026

TL;DR

  1. CF and XA both failed to establish reliable direct speech grounding under the current B2 recipes.

  2. Clean QA F1 remained around 0.11 for both routes, close to the B1 context-only baseline of 0.1075, while the matched text-question control reached 0.7730 F1.

  3. Wrong, zero, silent, or bypassed speech produced almost the same answers as correct speech, and Pair EM was 0/96 for both routes.

  4. Increasing ASR/QA training from R0 to R1 did not improve grounding, so insufficient training steps alone are unlikely to explain the failure.

  5. The remaining bottleneck is now localized to the direct speech-to-Qwen interface.

FastConformer features
↓
projector
↓
CF / XA
↓
Qwen

B3 should not begin until basic speech-dependent QA is established.

Previous Experiment

E003 M1b-B1 Grounding Task Validation established that the new answer-only paired-question task is valid.

It showed:

Context only
→ F1 = 0.1075
 
Context + text question
→ F1 = 0.7601
 
Speech
↓
FastConformer
↓
CTC transcript
↓
same QA model
→ F1 = 0.7338

Therefore:

  • Qwen3-0.6B can learn the QA task;

  • the task genuinely requires question information;

  • the frozen FastConformer preserves enough linguistic information for QA.

This narrowed the unresolved bottleneck to:

FastConformer speech representation
↓
projector + CF / XA
↓
Qwen

B2 directly tests whether this interface can establish speech grounding.

Question

Can CF and XA convert FastConformer speech representations into information that Qwen can use for question-dependent answer generation?

More specifically:

  1. Can correct spoken questions improve QA performance above the context-only baseline?

  2. Can the model distinguish two spoken questions that share the same context?

  3. Does changing or removing the speech change the generated answer?

  4. Is failure caused simply by insufficient ASR/QA training exposure?

  5. Are CF and XA ready for the generation-interference experiments in B3?

Hypothesis

If direct speech grounding is learned, then:

Context
+
Speech Question A
↓
Answer A
 
Context
+
Speech Question B
↓
Answer B

For paired questions sharing the same context:

swap speech
↓
answer should also switch

and removing the speech should cause a clear performance drop:

correct speech
>
wrong speech
>
zero / silence / bypass

The B2 development gate therefore requires both strong clean QA performance and clear causal dependence on speech content.

Description

B2 trains CF and XA directly on the B1 answer-only task.

Unlike B1:

text question
→ Qwen

B2 removes the question transcript from the prompt.

The only path through which Qwen can receive question information is:

Spoken question
↓
Frozen FastConformer
↓
speech features
↓
projector
↓
CF / XA
↓
Qwen
↓
Answer

The text context is still given directly to Qwen.

Shared Setup

Both routes use:

  • Qwen3-0.6B;

  • frozen FastConformer;

  • the same 2,048 training questions;

  • 192 tune questions / 96 context pairs;

  • answer-only QA supervision;

  • identical LoRA initialization;

  • identical projector initialization;

  • identical training order;

  • FP32 training;

  • no transcript input during QA inference.

CF and XA have approximately matched trainable parameter counts.

CF: 19.01M trainable parameters
XA: 19.30M trainable parameters

R0 — Baseline Recipe

300 ASR updates
+
600 QA updates

Both CF and XA are trained independently under the same exposure schedule.

R1 — Longer Alignment Training

Because R0 failed the development criteria, B2 activates R1:

1200 ASR updates
+
1200 QA updates

The purpose is to test whether the R0 failure was mainly caused by insufficient training exposure.

R1 increases both speech-transcription training and QA training while keeping the route architecture unchanged.

Speech Dependency Controls

Final checkpoints are evaluated with:

clean
native wrong speech
same-length wrong speech
projected zero
silence
route bypass

The critical question is not only whether clean F1 is high.

The model must behave differently when the spoken question changes.

Result

R0

Clean QA

CF clean F1 = 0.1093
XA clean F1 = 0.1134
 
CF Pair EM = 0 / 96
XA Pair EM = 0 / 96

For comparison:

B1 context-only = 0.1075
B1 text question = 0.7601

Therefore, providing the spoken question through CF/XA produces almost no improvement over giving Qwen no question at all.

Speech Dependency

CF:

clean               0.1093
native wrong        0.1110
same-length wrong   0.1093
projected zero      0.1128
silence             0.1107
bypass              0.1128

XA:

clean               0.1134
native wrong        0.1075
same-length wrong   0.1134
projected zero      0.1134
silence             0.1134
bypass              0.1134

The model behaves almost identically regardless of whether the speech is:

correct
wrong
zero
silent
absent

The answer therefore shows almost no causal dependence on speech content.

Answer Switching

When the spoken question is replaced by the paired alternative question:

CF switching rate = 0%
XA switching rate = 0%

A grounded model should change its answer when the question changes.

Neither route does so.

R0 therefore fails the B2 development gate.

R1

R1 substantially increases training exposure:

R0:
300 ASR + 600 QA
 
R1:
1200 ASR + 1200 QA

However, final QA performance remains essentially unchanged.

CF clean F1 = 0.1081
XA clean F1 = 0.1111
 
CF Pair EM = 0 / 96
XA Pair EM = 0 / 96

Speech dependency also remains weak.

CF:

clean               0.1081
native wrong        0.1193
same-length wrong   0.0953
projected zero      0.1015
silence             0.1035
bypass              0.1015

XA:

clean               0.1111
native wrong        0.1074
same-length wrong   0.1111
projected zero      0.1111
silence             0.1111
bypass              0.1111

Answer switching remains almost absent:

CF switching rate ≈ 1%
XA switching rate = 0%

Therefore, simply increasing the R0 training budget does not establish speech grounding.

Matched Text Control

The matched R1 text-question control reaches:

F1 = 0.7730
Pair EM = 33 / 96

The contrast is therefore approximately:

text question
→ F1 = 0.773
 
speech through CF/XA
→ F1 ≈ 0.11

The underlying QA task remains easily learnable when the question reaches Qwen in textual form.

The failure is specific to the direct speech-conditioning path.

ASR Diagnostic

The route models also fail to reliably transcribe the question during free ASR decoding.

Final diagnostic WER remains around:

CF ≈ 0.92
XA ≈ 1.00

despite longer ASR training.

At the same time, teacher-forced ASR loss decreases during training.

This suggests a possible shortcut:

correct transcript prefix
↓
Qwen learns how questions usually continue

without necessarily learning:

speech features
↓
actual spoken content

Teacher-forced loss reduction therefore cannot be treated as evidence of speech grounding.

Free decoding and causal speech interventions remain necessary.

Why R2 Was Not Run

R2 would mix ASR and QA training to test whether QA training destroys speech recognition learned during the ASR stage.

However, R1 does not show the required ASR-forgetting pattern.

The final ASR WER is not substantially worse than the earlier ASR-stage diagnostic.

Therefore, the evidence does not support:

ASR grounding was learned
↓
QA training caused catastrophic forgetting

Instead, the more likely problem is:

robust speech-to-Qwen grounding
was never established

R2 was therefore not activated.

Why XA Capacity Was Not Increased

XA width128/256 is only meaningful if:

CF succeeds
but
XA width64 fails

That would indicate a route-specific XA capacity problem.

Instead:

CF fails
XA fails

Therefore, the current evidence points to a shared speech-to-language alignment problem rather than an XA-specific bottleneck.

Increasing XA width would not resolve the scientific ambiguity.

Interpretation

B2 strengthens the bottleneck identified in B0 and B1.

The evidence now forms the following chain:

[E002]
 
Speech
↓
FastConformer
↓
✓ useful information exists
 
but
 
speech features
↓
CF / XA
↓
Qwen
↓
✗ unreliable grounding

Then:

[E003]
 
text context + text question
↓
Qwen
↓
✓ QA works
 
Speech
↓
FastConformer
↓
CTC transcript
↓
Qwen
↓
✓ QA mostly works

Now B2 shows:

text context
+
Speech
↓
FastConformer
↓
projector
↓
CF / XA
↓
Qwen
↓
✗ QA remains near context-only

This makes the remaining bottleneck increasingly specific:

FastConformer features
        ↓
speech-language interface
        ↓
projector + CF / XA
        ↓
Qwen semantic computation

The current experiment does not show that CF or XA are fundamentally ineffective routing mechanisms.

It shows that:

Under the current training objective, initialization, and exposure budget, neither route learns to make Qwen depend reliably on spoken question semantics.

Therefore, B2 cannot be used to compare:

CF vs XA

because neither model has passed the prerequisite of understanding the speech.

A route that ignores speech may appear robust to interference simply because the interference never affects generation.

Why B3 Is Blocked

B3 is intended to study incoming speech during ongoing generation.

For example:

Qwen generating answer
        +
new user speech arrives
        ↓
CF / XA
        ↓
effect on ongoing generation

This comparison is only meaningful if both routes already understand speech.

Otherwise:

speech interference has little effect

could mean:

the route is robust

or simply:

the route ignores speech

B2 cannot distinguish these interpretations because both CF and XA currently fail basic grounding.

Therefore:

B3 = blocked

No B3/B4, confirmation/test inference, backbone upgrade, or M2 experiment should begin yet.

Conclusion

B2 rules out one simple explanation for the previous failure:

R0 was just too short

Increasing the training budget substantially does not improve grounding.

Together with B1, the current evidence suggests:

Qwen QA ability          ✓
FastConformer information ✓
Direct speech → Qwen      ✗

The next stage should therefore focus specifically on speech-language interface learnability and alignment, not on CF/XA scientific comparison.

Next Experiment

B2 Follow-up — Speech-to-Qwen Alignment Diagnosis

Step 1 — Tiny Causal Grounding Test

Before scaling anything, test whether the existing interface can overfit a very small dataset.

Use approximately:

8–16 context pairs
16–32 spoken questions

with both CF and XA.

Start from the validated B1 QA state and freeze Qwen.

Train only:

projector
+
CF / XA
+
necessary listening parameters

using the ordinary answer QA objective.

Success must require more than low training loss.

The model should satisfy:

Speech A → Answer A
Speech B → Answer B
 
swap A/B
→ answers also swap
 
zero / silence / bypass
→ performance collapses

and free ASR decoding should reflect actual speech content rather than only teacher-forced NLL improvement.

Decision

If the tiny set succeeds:

architecture is learnable
↓
investigate generalization / exposure / curriculum

If the tiny set fails:

basic interface learning is broken
↓
inspect representation and optimization

including:

  • projector output scale vs Qwen token embeddings;

  • projector initialization;

  • CF speech contribution magnitude;

  • XA gate and residual contribution;

  • route parameter update magnitude;

  • correct-vs-wrong speech logit differences;

  • speech-conditioned vs zero-conditioned hidden states.

Step 2 — Stronger Alignment Supervision

If Frozen-Qwen + ordinary QA CE still fails, test stronger objectives separately.

First candidate:

Text Teacher
context + text question
↓
validated B1 Qwen
↓
next-token distribution
        │
        │ KD
        ▼
Speech Student
context + speech
↓
FastConformer
↓
projector + CF / XA
↓
same frozen Qwen

Train with:

QA loss
+
speech/text behavioral distillation

The transcript may be used by the teacher during training but must never enter the speech student’s inference input.

If this still fails, separately test a more direct auxiliary speech-semantic objective such as:

CTC-style transcript alignment

or:

speech-conditioned hidden representation
≈
text-conditioned hidden representation

Each intervention should be tested independently so that the cause of any improvement remains identifiable.

Do Not Scale Yet

Until tiny causal grounding succeeds, do not:

Qwen3-0.6B → Qwen3-1.7B
 
XA width64 → width128/256
 
massively increase training steps
 
enter B3

The immediate goal is simpler:

Make Qwen’s answer causally depend on what was spoken.