CompleteDate: Sep 13, 2026

TL;DR

  1. E0 tested whether the existing CF/XA speech-to-Qwen interfaces can at least overfit speech semantics on a tiny controlled dataset.

  2. CF showed the first clear evidence of speech-content dependence: QA EM reached 0.844, candidate switching reached 0.750, and clean speech outperformed projected-zero speech by +0.310 F1 on fitted examples.

  3. However, CF still failed 3 of 4 preregistered QA fit targets and generalized poorly to the train-derived diagnostic subset.

  4. XA did not demonstrate semantic grounding. On all 64 fitted and 16 diagnostic examples, clean outputs could be reproduced when speech content was removed.

  5. Independent ASR fitting also failed: WER remained 0.513 for CF and 0.852 for XA.

  6. The evidence therefore suggests that the main remaining bottleneck is the interface:

    FastConformer representation
    ↓
    projector + CF / XA
    ↓
    representation usable by Qwen
  7. The next experiment is E1: start from the already capable B1 text-QA model, freeze Qwen and its language LoRA, and train only the speech-side interface.

B3 remains blocked.

Previous Experiment

E004 Routing Fail Translating showed that direct speech grounding failed under both the R0 and longer R1 training recipes.

The main result was:

Text question
↓
Qwen
→ F1 ≈ 0.77
 
Speech question
↓
FastConformer
↓
projector + CF / XA
↓
Qwen
→ F1 ≈ 0.11

Correct, wrong, zero, silent, and bypassed speech produced almost identical answers.

Increasing training from:

R0:
300 ASR + 600 QA

to:

R1:
1200 ASR + 1200 QA

did not solve the problem.

This localized the remaining bottleneck to:

FastConformer features
↓
speech-to-Qwen interface
↓
Qwen

However, B2 could not distinguish two possibilities:

A. The interface is fundamentally unable to learn speech grounding
 
or
 
B. The full-data optimization problem is too difficult,
   even though the interface is locally learnable

E0 was designed to separate these possibilities.

Question

Can the existing CF and XA interfaces learn genuine speech-dependent behavior when asked to overfit a very small controlled dataset?

More specifically:

  1. Can the model fit the spoken QA examples?

  2. Does changing the speech content change the answer?

  3. Does removing speech information hurt performance?

  4. Can the interface independently learn transcription from speech?

  5. Is the B2 failure caused by complete interface failure, or by optimization/generalization difficulty?

Hypothesis

If the interface is fundamentally learnable, then on a sufficiently small fitted dataset:

Speech Question A
↓
route
↓
Answer A
 
Speech Question B
↓
route
↓
Answer B

Changing the speech should change the answer:

correct speech
>
wrong speech
>
zero / bypass

For paired questions sharing the same context:

Speech A → Answer A
Speech B → Answer B
 
swap A ↔ B
→
answer should also switch

A model that merely memorizes context or timing without using speech semantics should fail these causal controls.

Description

E0 is a tiny controlled fitting experiment rather than a generalization experiment.

It uses:

32 paragraph pairs
64 fitted questions
 
+
 
8 additional paragraph pairs
16 diagnostic questions

The diagnostic examples are disjoint from the E0 fitting examples at the paragraph/article level, but both sets come from the global training split.

Therefore:

E0 measures local interface learnability, not tune-set generalization.

Four Independent Fits

CF and XA are each trained separately for:

ASR fitting
QA fitting

producing four independent runs:

CF-ASR
XA-ASR
CF-QA
XA-QA

Each receives exactly:

300 updates

for a total of:

1200 updates

ASR and QA use separate optimizers and checkpoints.

Initialization

E0 reconstructs the original B2 initialization:

Qwen3-0.6B
+
fresh language LoRA
+
fresh projector
+
fresh WAIT parameters
+
fresh CF / XA route

It does not initialize from the trained B1 text model.

QA Speech Controls

At the final checkpoint, the fitted and diagnostic questions are evaluated under:

clean speech
native wrong speech
same-length wrong speech
projected zero
route bypass

The important question is not simply:

Can the model produce the right answer?

but:

Does the answer causally depend on speech content?

E0 QA Fit Targets

The preregistered tiny-fit targets are:

EM ≥ 0.90
 
Pair EM ≥ 0.80
 
candidate correct switching ≥ 0.80
 
clean − zero F1 ≥ 0.20

Independent ASR fitting requires:

WER ≤ 0.10

These are technical E0 diagnosis targets, not B3 gates.

Result

CF QA Fit

CF shows substantial improvement compared with B2.

Clean F1            = 0.844
EM                  = 0.844
Pair EM             = 0.688
Candidate switching = 0.750
Clean − zero F1     = +0.310

The speech interventions show clear content dependence:

clean − native wrong      ≈ +0.683
clean − same-length wrong ≈ +0.466
clean − projected zero    ≈ +0.310

This is the first strong evidence that the CF path can actually make Qwen depend on spoken content.

The important result is therefore not simply:

CF F1 = 0.844

but:

correct speech
↓
different model behavior
 
wrong / removed speech
↓
performance decreases

Therefore:

CF is not fundamentally disconnected from speech semantics.

However, CF still fails the full E0 target.

It passes only:

clean − zero F1 ≥ 0.20

while missing:

EM ≥ 0.90
Pair EM ≥ 0.80
candidate switching ≥ 0.80

So CF demonstrates partial learnability, not successful grounding.

CF Diagnostic Generalization

Performance drops sharply on the 16 train-derived diagnostic questions:

Clean F1       = 0.219
EM             = 0.188
Pair EM        = 0
Candidate switch = 0
Clean − zero F1 ≈ +0.083

Therefore:

tiny fitted examples
→ some real speech dependence
 
new examples
→ weak grounding

E0 does not establish generalizable CF grounding.

XA QA Fit

XA appears to partially fit the QA task if only ordinary QA accuracy is examined:

Clean F1            = 0.581
EM                  = 0.578
Pair EM             = 0.156
Candidate switching = 0.375

However:

Clean − zero F1 = 0.000

and the causal controls reveal a much more important result.

For all:

64 fitted questions
+
16 diagnostic questions

the XA clean output can be reproduced when speech content is removed.

In particular:

clean
same-length wrong
projected zero
bypass

produce the same outputs across the evaluated cases.

Therefore, XA’s apparent QA fitting cannot be interpreted as semantic speech grounding.

A plausible behavior is instead:

context
+
WAIT / timeline structure
+
language-model prior
↓
answer

with little meaningful contribution from:

speech content

Thus:

XA does not demonstrate semantic grounding in E0.

XA Signal Magnitude

The numerical diagnostics support this interpretation.

At the final XA-QA checkpoint, the contribution of the speech-attention path to the hidden state is extremely small.

Observed layer attention residual / hidden RMS is approximately:

7.5e-9
to
5.8e-6

and replacing speech with same-length alternative speech changes logits by only approximately:

8e-6

The parameters and gradients are not zero, so this is not evidence of a completely disconnected optimizer.

Instead:

XA parameters update
↓
but speech information has negligible effect
↓
Qwen output remains nearly speech-independent

The exact cause remains unresolved.

Possible explanations include:

  • weak interface optimization;

  • representation-scale mismatch;

  • context/timing shortcuts;

  • ineffective XA contribution;

  • poor speech-language alignment.

E0 does not yet distinguish these causes.

Independent ASR Fit

E0 also independently asks each route to transcribe the speech.

Final fitted-set WER:

CF WER = 0.513
XA WER = 0.852

Target:

WER ≤ 0.10

Therefore, both fail.

CF again performs substantially better than XA, but neither establishes reliable speech-to-language mapping.

The diagnostic WER is even worse:

CF diagnostic WER ≈ 0.968
XA diagnostic WER ≈ 0.916

This is consistent with the QA result:

CF
→ some local speech-content learning
 
XA
→ much weaker semantic speech use
 
both
→ poor generalization

Teacher-Forced Loss Is Still Misleading

All four runs substantially reduce training cross-entropy.

For example:

CF-ASR:
CE ≈ 4.07 → 0.27
 
XA-ASR:
CE ≈ 4.09 → 0.37
 
CF-QA:
CE ≈ 2.17 → 0.12
 
XA-QA:
CE ≈ 2.18 → 0.17

Yet free ASR decoding and speech-dependent QA remain poor.

Therefore:

low teacher-forced CE
≠
speech grounding

The model can learn strong next-token prediction under reference prefixes without learning a representation robust enough for free generation.

This reinforces the need for:

free decoding
+
speech intervention controls

rather than training loss alone.

Numerical / Engineering Checks

The failure is not explained by an obvious broken training pipeline.

Checks confirm that:

projector parameters update
LoRA parameters update
route parameters update
gradients exist
frozen parameters remain frozen
labels are correctly shifted
cache/reference behavior is valid
checkpoints remain finite

A supplemental numerical check initially failed because a helper compared sequences with different EOS-appended input shapes.

The failure was preserved rather than rewritten as a pass.

A corrected version of the numerical test separated:

same-input label validation

from:

changed-shape prefix-logit invariance

and passed without changing any trained model, optimizer state, scientific threshold, or primary result.

Therefore, E0 does not currently indicate a basic implementation or numerical correctness failure.

Interpretation

E0 changes the interpretation of the B2 failure.

Before E0, the situation was approximately:

FastConformer
↓
projector + CF / XA
↓
Qwen
↓
grounding fails
 
Unknown:
Can this interface learn speech semantics at all?

After E0:

FastConformer
↓
speech representation
↓
projector + CF
↓
Qwen
↓
partial speech-content dependence
✓ locally learnable
 
FastConformer
↓
speech representation
↓
projector + XA
↓
Qwen
↓
almost no semantic speech dependence
✗ grounding not demonstrated

This means the shared B2 failure cannot simply be described as:

speech never reaches Qwen

For CF, speech can influence Qwen semantically under tiny fitted conditions.

The more precise bottleneck is:

FastConformer representation
↓
speech-to-Qwen interface
↓
stable semantic representation
↓
generalizable Qwen behavior

CF can partially learn this mapping.

XA currently shows almost no convincing evidence that it can.

What E0 Rules Out

1. CF Is Not Completely Broken

The clean-vs-zero and clean-vs-wrong differences show that CF can make the answer depend on speech content.

Therefore:

CF speech path is completely disconnected

is unlikely.

2. More Training Steps Alone Are Not the Main Answer

E004 already showed that increasing full-data training from R0 to R1 did not solve grounding.

E0 now shows that even strong fitting behavior does not automatically generalize.

Therefore the problem is not simply:

train longer

3. Training Loss Cannot Be Used as Grounding Evidence

Strong loss reduction occurs even when:

free ASR fails

or:

speech removal does not change the answer

4. The Current Failure Is Not Yet Evidence That CF Is Scientifically Better Than XA

E0 suggests that CF is easier to optimize under the current setup.

However:

CF = partially grounded
XA = not grounded

does not yet support a meaningful routing comparison.

Comparing interference robustness requires both routes to first understand speech.

Why B3 Is Still Blocked

B3 aims to compare how CF and XA handle incoming speech during ongoing generation.

That experiment assumes:

normal speech
→ both routes understand it

before testing:

interference
interruptions
wrong speaker
uncertain speech

Currently:

CF
→ partial tiny-fit grounding
→ poor diagnostic generalization
 
XA
→ no convincing semantic grounding

Therefore, a B3 result could not distinguish:

route is robust to interference

from:

route simply ignores speech

B3 remains blocked.

Before B3, the project still requires:

  • a common full-data tune grounding pass;

  • seed42 replication;

  • generation-interference task validation;

  • positive controls;

  • data/reference review;

  • explicit route matching and capacity disclosure.

Conclusion

E0 provides the first positive evidence that the direct speech-to-Qwen architecture is not completely hopeless.

The current picture is:

FastConformer speech information
✓ available
 
Qwen QA ability
✓ available
 
CF interface
△ locally learnable but weak / non-generalizing
 
XA interface
✗ semantic grounding not demonstrated

The central problem is therefore increasingly specific:

How can FastConformer speech representations be converted into a stable representation that an already capable Qwen language model can reliably use?

The next experiment should isolate this interface from simultaneous language-model adaptation.

Next Experiment

E1 — Frozen-Language Interface Diagnosis

E1 starts from the already successful B1 text-QA model.

Instead of jointly changing the speech interface and language model:

speech interface
        ↓
Qwen + LoRA also changing
        ↓
answer

E1 freezes the language side:

FastConformer
     🔒
      ↓
speech features
      ↓
projector
+
CF / XA
+
WAIT
      ↑
   train
      ↓
B1 Qwen + language LoRA
     🔒
      ↓
answer

The key idea is:

Qwen already knows how to solve the QA task. Keep its representation space fixed and force the speech interface to learn how to communicate with it.

Initialization

Use:

B1 text model at step 600

Freeze:

Qwen base
+
B1 language LoRA

Train fresh:

projector
+
CF / XA
+
WAIT / listening parameters

Training

For each route:

1200 ASR updates
+
1200 QA updates

Unlike E0, E1 uses the full training exposure and evaluates the complete tune grounding gate.

Key Question

Can the speech interface learn grounding when the target language model representation is stable?

Possible outcomes:

CF succeeds
XA succeeds
↓
joint language/interface optimization was likely a major B2 problem
↓
replicate with seed42

or:

CF succeeds
XA fails
↓
investigate XA-specific optimization / capacity / routing behavior

or:

CF fails
XA fails
↓
ordinary QA/ASR CE is insufficient
↓
move to stronger alignment supervision

such as the preregistered E2 CTC-style auxiliary alignment.

Immediate Goal

Do not enter B3 yet.

The immediate objective remains:

Make Qwen’s output reliably and causally depend on what was spoken, on unseen examples.