TL;DR
E0 tested whether the existing CF/XA speech-to-Qwen interfaces can at least overfit speech semantics on a tiny controlled dataset.
CF showed the first clear evidence of speech-content dependence: QA EM reached 0.844, candidate switching reached 0.750, and clean speech outperformed projected-zero speech by +0.310 F1 on fitted examples.
However, CF still failed 3 of 4 preregistered QA fit targets and generalized poorly to the train-derived diagnostic subset.
XA did not demonstrate semantic grounding. On all 64 fitted and 16 diagnostic examples, clean outputs could be reproduced when speech content was removed.
Independent ASR fitting also failed: WER remained 0.513 for CF and 0.852 for XA.
The evidence therefore suggests that the main remaining bottleneck is the interface:
FastConformer representation ↓ projector + CF / XA ↓ representation usable by QwenThe next experiment is E1: start from the already capable B1 text-QA model, freeze Qwen and its language LoRA, and train only the speech-side interface.
B3 remains blocked.
Previous Experiment
E004 Routing Fail Translating showed that direct speech grounding failed under both the R0 and longer R1 training recipes.
The main result was:
Text question
↓
Qwen
→ F1 ≈ 0.77
Speech question
↓
FastConformer
↓
projector + CF / XA
↓
Qwen
→ F1 ≈ 0.11Correct, wrong, zero, silent, and bypassed speech produced almost identical answers.
Increasing training from:
R0:
300 ASR + 600 QAto:
R1:
1200 ASR + 1200 QAdid not solve the problem.
This localized the remaining bottleneck to:
FastConformer features
↓
speech-to-Qwen interface
↓
QwenHowever, B2 could not distinguish two possibilities:
A. The interface is fundamentally unable to learn speech grounding
or
B. The full-data optimization problem is too difficult,
even though the interface is locally learnableE0 was designed to separate these possibilities.
Question
Can the existing CF and XA interfaces learn genuine speech-dependent behavior when asked to overfit a very small controlled dataset?
More specifically:
-
Can the model fit the spoken QA examples?
-
Does changing the speech content change the answer?
-
Does removing speech information hurt performance?
-
Can the interface independently learn transcription from speech?
-
Is the B2 failure caused by complete interface failure, or by optimization/generalization difficulty?
Hypothesis
If the interface is fundamentally learnable, then on a sufficiently small fitted dataset:
Speech Question A
↓
route
↓
Answer A
Speech Question B
↓
route
↓
Answer BChanging the speech should change the answer:
correct speech
>
wrong speech
>
zero / bypassFor paired questions sharing the same context:
Speech A → Answer A
Speech B → Answer B
swap A ↔ B
→
answer should also switchA model that merely memorizes context or timing without using speech semantics should fail these causal controls.
Description
E0 is a tiny controlled fitting experiment rather than a generalization experiment.
It uses:
32 paragraph pairs
64 fitted questions
+
8 additional paragraph pairs
16 diagnostic questionsThe diagnostic examples are disjoint from the E0 fitting examples at the paragraph/article level, but both sets come from the global training split.
Therefore:
E0 measures local interface learnability, not tune-set generalization.
Four Independent Fits
CF and XA are each trained separately for:
ASR fitting
QA fittingproducing four independent runs:
CF-ASR
XA-ASR
CF-QA
XA-QAEach receives exactly:
300 updatesfor a total of:
1200 updatesASR and QA use separate optimizers and checkpoints.
Initialization
E0 reconstructs the original B2 initialization:
Qwen3-0.6B
+
fresh language LoRA
+
fresh projector
+
fresh WAIT parameters
+
fresh CF / XA routeIt does not initialize from the trained B1 text model.
QA Speech Controls
At the final checkpoint, the fitted and diagnostic questions are evaluated under:
clean speech
native wrong speech
same-length wrong speech
projected zero
route bypassThe important question is not simply:
Can the model produce the right answer?but:
Does the answer causally depend on speech content?E0 QA Fit Targets
The preregistered tiny-fit targets are:
EM ≥ 0.90
Pair EM ≥ 0.80
candidate correct switching ≥ 0.80
clean − zero F1 ≥ 0.20Independent ASR fitting requires:
WER ≤ 0.10These are technical E0 diagnosis targets, not B3 gates.
Result
CF QA Fit
CF shows substantial improvement compared with B2.
Clean F1 = 0.844
EM = 0.844
Pair EM = 0.688
Candidate switching = 0.750
Clean − zero F1 = +0.310The speech interventions show clear content dependence:
clean − native wrong ≈ +0.683
clean − same-length wrong ≈ +0.466
clean − projected zero ≈ +0.310This is the first strong evidence that the CF path can actually make Qwen depend on spoken content.
The important result is therefore not simply:
CF F1 = 0.844but:
correct speech
↓
different model behavior
wrong / removed speech
↓
performance decreasesTherefore:
CF is not fundamentally disconnected from speech semantics.
However, CF still fails the full E0 target.
It passes only:
clean − zero F1 ≥ 0.20while missing:
EM ≥ 0.90
Pair EM ≥ 0.80
candidate switching ≥ 0.80So CF demonstrates partial learnability, not successful grounding.
CF Diagnostic Generalization
Performance drops sharply on the 16 train-derived diagnostic questions:
Clean F1 = 0.219
EM = 0.188
Pair EM = 0
Candidate switch = 0
Clean − zero F1 ≈ +0.083Therefore:
tiny fitted examples
→ some real speech dependence
new examples
→ weak groundingE0 does not establish generalizable CF grounding.
XA QA Fit
XA appears to partially fit the QA task if only ordinary QA accuracy is examined:
Clean F1 = 0.581
EM = 0.578
Pair EM = 0.156
Candidate switching = 0.375However:
Clean − zero F1 = 0.000and the causal controls reveal a much more important result.
For all:
64 fitted questions
+
16 diagnostic questionsthe XA clean output can be reproduced when speech content is removed.
In particular:
clean
same-length wrong
projected zero
bypassproduce the same outputs across the evaluated cases.
Therefore, XA’s apparent QA fitting cannot be interpreted as semantic speech grounding.
A plausible behavior is instead:
context
+
WAIT / timeline structure
+
language-model prior
↓
answerwith little meaningful contribution from:
speech contentThus:
XA does not demonstrate semantic grounding in E0.
XA Signal Magnitude
The numerical diagnostics support this interpretation.
At the final XA-QA checkpoint, the contribution of the speech-attention path to the hidden state is extremely small.
Observed layer attention residual / hidden RMS is approximately:
7.5e-9
to
5.8e-6and replacing speech with same-length alternative speech changes logits by only approximately:
8e-6The parameters and gradients are not zero, so this is not evidence of a completely disconnected optimizer.
Instead:
XA parameters update
↓
but speech information has negligible effect
↓
Qwen output remains nearly speech-independentThe exact cause remains unresolved.
Possible explanations include:
-
weak interface optimization;
-
representation-scale mismatch;
-
context/timing shortcuts;
-
ineffective XA contribution;
-
poor speech-language alignment.
E0 does not yet distinguish these causes.
Independent ASR Fit
E0 also independently asks each route to transcribe the speech.
Final fitted-set WER:
CF WER = 0.513
XA WER = 0.852Target:
WER ≤ 0.10Therefore, both fail.
CF again performs substantially better than XA, but neither establishes reliable speech-to-language mapping.
The diagnostic WER is even worse:
CF diagnostic WER ≈ 0.968
XA diagnostic WER ≈ 0.916This is consistent with the QA result:
CF
→ some local speech-content learning
XA
→ much weaker semantic speech use
both
→ poor generalizationTeacher-Forced Loss Is Still Misleading
All four runs substantially reduce training cross-entropy.
For example:
CF-ASR:
CE ≈ 4.07 → 0.27
XA-ASR:
CE ≈ 4.09 → 0.37
CF-QA:
CE ≈ 2.17 → 0.12
XA-QA:
CE ≈ 2.18 → 0.17Yet free ASR decoding and speech-dependent QA remain poor.
Therefore:
low teacher-forced CE
≠
speech groundingThe model can learn strong next-token prediction under reference prefixes without learning a representation robust enough for free generation.
This reinforces the need for:
free decoding
+
speech intervention controlsrather than training loss alone.
Numerical / Engineering Checks
The failure is not explained by an obvious broken training pipeline.
Checks confirm that:
projector parameters update
LoRA parameters update
route parameters update
gradients exist
frozen parameters remain frozen
labels are correctly shifted
cache/reference behavior is valid
checkpoints remain finiteA supplemental numerical check initially failed because a helper compared sequences with different EOS-appended input shapes.
The failure was preserved rather than rewritten as a pass.
A corrected version of the numerical test separated:
same-input label validationfrom:
changed-shape prefix-logit invarianceand passed without changing any trained model, optimizer state, scientific threshold, or primary result.
Therefore, E0 does not currently indicate a basic implementation or numerical correctness failure.
Interpretation
E0 changes the interpretation of the B2 failure.
Before E0, the situation was approximately:
FastConformer
↓
projector + CF / XA
↓
Qwen
↓
grounding fails
Unknown:
Can this interface learn speech semantics at all?After E0:
FastConformer
↓
speech representation
↓
projector + CF
↓
Qwen
↓
partial speech-content dependence
✓ locally learnable
FastConformer
↓
speech representation
↓
projector + XA
↓
Qwen
↓
almost no semantic speech dependence
✗ grounding not demonstratedThis means the shared B2 failure cannot simply be described as:
speech never reaches QwenFor CF, speech can influence Qwen semantically under tiny fitted conditions.
The more precise bottleneck is:
FastConformer representation
↓
speech-to-Qwen interface
↓
stable semantic representation
↓
generalizable Qwen behaviorCF can partially learn this mapping.
XA currently shows almost no convincing evidence that it can.
What E0 Rules Out
1. CF Is Not Completely Broken
The clean-vs-zero and clean-vs-wrong differences show that CF can make the answer depend on speech content.
Therefore:
CF speech path is completely disconnectedis unlikely.
2. More Training Steps Alone Are Not the Main Answer
E004 already showed that increasing full-data training from R0 to R1 did not solve grounding.
E0 now shows that even strong fitting behavior does not automatically generalize.
Therefore the problem is not simply:
train longer3. Training Loss Cannot Be Used as Grounding Evidence
Strong loss reduction occurs even when:
free ASR failsor:
speech removal does not change the answer4. The Current Failure Is Not Yet Evidence That CF Is Scientifically Better Than XA
E0 suggests that CF is easier to optimize under the current setup.
However:
CF = partially grounded
XA = not groundeddoes not yet support a meaningful routing comparison.
Comparing interference robustness requires both routes to first understand speech.
Why B3 Is Still Blocked
B3 aims to compare how CF and XA handle incoming speech during ongoing generation.
That experiment assumes:
normal speech
→ both routes understand itbefore testing:
interference
interruptions
wrong speaker
uncertain speechCurrently:
CF
→ partial tiny-fit grounding
→ poor diagnostic generalization
XA
→ no convincing semantic groundingTherefore, a B3 result could not distinguish:
route is robust to interferencefrom:
route simply ignores speechB3 remains blocked.
Before B3, the project still requires:
-
a common full-data tune grounding pass;
-
seed42 replication;
-
generation-interference task validation;
-
positive controls;
-
data/reference review;
-
explicit route matching and capacity disclosure.
Conclusion
E0 provides the first positive evidence that the direct speech-to-Qwen architecture is not completely hopeless.
The current picture is:
FastConformer speech information
✓ available
Qwen QA ability
✓ available
CF interface
△ locally learnable but weak / non-generalizing
XA interface
✗ semantic grounding not demonstratedThe central problem is therefore increasingly specific:
How can FastConformer speech representations be converted into a stable representation that an already capable Qwen language model can reliably use?
The next experiment should isolate this interface from simultaneous language-model adaptation.
Next Experiment
E1 — Frozen-Language Interface Diagnosis
E1 starts from the already successful B1 text-QA model.
Instead of jointly changing the speech interface and language model:
speech interface
↓
Qwen + LoRA also changing
↓
answerE1 freezes the language side:
FastConformer
🔒
↓
speech features
↓
projector
+
CF / XA
+
WAIT
↑
train
↓
B1 Qwen + language LoRA
🔒
↓
answerThe key idea is:
Qwen already knows how to solve the QA task. Keep its representation space fixed and force the speech interface to learn how to communicate with it.
Initialization
Use:
B1 text model at step 600Freeze:
Qwen base
+
B1 language LoRATrain fresh:
projector
+
CF / XA
+
WAIT / listening parametersTraining
For each route:
1200 ASR updates
+
1200 QA updatesUnlike E0, E1 uses the full training exposure and evaluates the complete tune grounding gate.
Key Question
Can the speech interface learn grounding when the target language model representation is stable?
Possible outcomes:
CF succeeds
XA succeeds
↓
joint language/interface optimization was likely a major B2 problem
↓
replicate with seed42or:
CF succeeds
XA fails
↓
investigate XA-specific optimization / capacity / routing behavioror:
CF fails
XA fails
↓
ordinary QA/ASR CE is insufficient
↓
move to stronger alignment supervisionsuch as the preregistered E2 CTC-style auxiliary alignment.
Immediate Goal
Do not enter B3 yet.
The immediate objective remains:
Make Qwen’s output reliably and causally depend on what was spoken, on unseen examples.