CompleteDate: Sep 15, 2026

TL;DR

  1. E1 tested whether the speech interface can learn grounding when the already capable B1 text-QA model is frozen, removing simultaneous language-model adaptation as a possible source of failure.

  2. CF showed limited but statistically supported speech-content dependence on tune: clean speech outperformed same-length wrong speech by +0.0785 F1, with 95% CI [0.0320, 0.1284].

  3. However, CF still failed semantic grounding: clean F1 was only 0.155, Pair EM was 0, candidate switching was 4.17%, and free correct switching was 0.

  4. XA showed almost no semantic dependence on speech content. Replacing clean speech with same-length wrong speech left 188/192 outputs unchanged, and silence left 183/192 unchanged.

  5. Both routes also showed poor free ASR after the ASR stage, indicating that the speech-to-language mapping itself remains weak.

  6. Since the frozen B1 text model reaches approximately 0.76 tune F1, while speech-conditioned CF/XA remain near 0.10–0.15, the bottleneck is increasingly localized to:

    FastConformer representation
    ↓
    projector + routing interface
    ↓
    representation usable by Qwen
  7. E1 therefore rules out the hypothesis that simply freezing a capable language model is sufficient. The remaining problem is more specifically speech-language alignment / interface learning.

  8. The next preregistered experiment is E2, which adds direct-projector CTC supervision to explicitly align projected speech representations with Qwen’s lexical space.

B3 remains blocked.

Previous Experiment

E005 Tiny Overfit Testing asked whether CF and XA could at least learn speech semantics on a tiny fitted dataset.

E0 found:

CF
→ strong fitted QA performance
→ clear speech-content dependence
→ poor generalization
 
XA
→ apparent QA fitting
→ almost no genuine speech-content dependence

For CF:

Fit F1              = 0.844
Pair EM             = 0.688
Candidate switching = 0.750
Clean − zero F1     = +0.310

This showed that CF was not fundamentally disconnected from speech.

However, performance dropped sharply on diagnostic questions:

Diagnostic F1       = 0.219
Pair EM             = 0
Candidate switching = 0

XA showed even weaker evidence of semantic grounding.

The key remaining ambiguity was:

Maybe the interface is learnable,
but jointly adapting the speech interface
and language model makes optimization unstable.
 

E1 therefore fixes the language side and asks the speech interface to adapt to a stable target representation.

Question

Can CF and XA learn reliable speech grounding when the language model is already capable of the QA task and remains frozen?

More specifically:

  1. Can a fresh speech interface learn to communicate with a fixed Qwen representation space?

  2. Does correct speech produce better answers than wrong or removed speech?

  3. Does changing the spoken question cause the answer to switch correctly?

  4. Was simultaneous language-model adaptation a major cause of the previous failures?

  5. Can both routes establish sufficient grounding to proceed toward B3?

Hypothesis

B1 already established that the language model can solve the QA task from text:

Text Question
↓
Qwen
→ tune F1 ≈ 0.76

Therefore E1 freezes:

Qwen base
+
B1 language LoRA

and trains only:

FastConformer
🔒
↓
speech features
↓
projector
+
CF / XA
+
WAIT
↓
frozen Qwen

If the main problem in previous experiments was simultaneous language/interface adaptation, then fixing the target language space should make speech grounding easier.

Expected behavior:

correct speech
↓
correct semantic interpretation
↓
correct answer
 
wrong speech
↓
different interpretation
↓
different answer

A successful interface should therefore show both:

clean > wrong / zero

and:

spoken question changes
↓
answer correctly switches

Description

Both routes start from the same B1 text-QA model at:

step 600

Freeze:

Qwen base
+
B1 language LoRA

Freshly initialize and train:

projector
+
CF / XA
+
WAIT

Each route receives:

1200 ASR updates
+
1200 QA updates

The optimizer state is carried from ASR into QA, while the learning-rate schedule is reset for the QA stage.

Speech Controls

The final models are evaluated under:

clean
 
native wrong
 
same-length wrong
 
projected zero
 
silence
 
bypass

The important distinction is:

speech path affects model

versus:

speech CONTENT affects model

For example:

clean vs projected zero

tests whether the speech pathway matters.

But:

clean vs same-length wrong

is stronger evidence of semantic dependence because timing and approximate length are preserved while speech content changes.

Grounding Gate

For each route, E1 requires:

Clean F1 ≥ 0.30
 
Pair EM ≥ 0.20
 
Candidate switching ≥ 0.50

and:

clean − native wrong F1 ≥ 0.05
 
clean − projected zero F1 ≥ 0.05

with positive paired bootstrap confidence intervals.

Free-answer switching must also support the candidate-based results.

Neither route passed the complete gate.

Result

CF

Tune performance:

Clean F1              = 0.155
EM                    = 0.109
Pair EM               = 0
Candidate switching   = 4.17%
Free correct switching = 0

CF therefore remains far below the required grounding level.

However, unlike previous full-data experiments, CF now shows clear evidence that speech content itself affects generation.

Important controls:

clean − native wrong
= +0.0716
 
clean − same-length wrong
= +0.0785
 
clean − projected zero
= +0.0797
 
clean − silence
= +0.0629

For the particularly important same-length wrong control:

ΔF1 = +0.0785
 
95% CI
= [0.0320, 0.1284]

The entire interval is above zero.

Therefore:

correct speech content
↓
improves CF output

even when timing / approximate speech length are controlled.

This is stronger evidence than simply observing a clean-vs-zero difference.

What CF Has Learned

CF appears to have reached:

speech content
↓
changes model behavior

but not:

speech semantics
↓
correctly determine answer

The distinction is important.

The model is sensitive to what was spoken, but that information is not yet decoded accurately enough to produce reliable semantic switching.

Thus:

CF demonstrates weak speech-content dependence, not complete semantic grounding.

XA

Tune performance:

Clean F1              = 0.098
EM                    = 0.073
Pair EM               = 0
Candidate switching   = 2.08%
Free correct switching = 0

More importantly, XA shows almost no sensitivity to speech content.

For:

clean
vs
same-length wrong

the generated token sequence remained identical on:

188 / 192 examples

For:

clean
vs
silence

it remained identical on:

183 / 192 examples

The clean-vs-control F1 differences were approximately:

clean − native wrong
≈ -0.001
 
clean − same-length wrong
≈ +0.001
 
clean − projected zero
≈ -0.014

Therefore:

change speech content
↓
almost no change in output

XA does not demonstrate semantic grounding.

XA Bypass Interpretation

Some XA outputs do change when the adapter is completely bypassed or zeroed.

However, this is not evidence that XA understands speech.

There is an important difference between:

change speech content

and:

remove / alter the whole speech adapter

The latter can change Qwen’s hidden-state dynamics even if the model does not use lexical speech information.

Therefore:

bypass changes output

only demonstrates that the XA module can influence the network.

It does not demonstrate:

speech semantics
↓
correctly influence generation

The same-length wrong control is much more informative, and XA essentially fails it.

ASR Diagnosis

The ASR stage also fails to establish reliable free transcription.

At ASR step 1200:

CF train WER ≈ 1.03
XA train WER ≈ 1.22

Both are extremely poor.

This suggests that even before QA:

speech
↓
linguistic representation usable by Qwen

has not been reliably established.

This is consistent with the QA results.

For CF:

some speech-content information survives

but not enough to reconstruct semantics reliably.

For XA:

speech-content information has almost no observable effect

Teacher-Forced Loss Remains Misleading

Training CE decreases substantially for both routes.

For example:

CF QA:
≈ 3.41 → 1.06
 
XA QA:
≈ 3.43 → 1.21

Yet free QA grounding remains poor.

Similarly, ASR CE decreases while free transcription remains extremely weak.

Therefore E1 reinforces:

low teacher-forced CE
≠
successful speech grounding

A model can improve next-token prediction under gold prefixes without learning a representation that works during free generation.

Grounding must therefore continue to be judged using:

free generation
 
+
 
causal speech controls

rather than loss curves alone.

Numerical / Engineering Checks

E1 does not reveal an obvious broken implementation.

The checks confirm:

projector / route parameters update
 
speech-path gradients exist
 
frozen Qwen remains frozen
 
cache/reference behavior is valid
 
causal timing checks pass
 
checkpoints remain finite
 
CF/XA interfaces are active

All required initial/final route checks pass.

Therefore the scientific failure cannot currently be explained by:

no gradients
 
wrongly frozen interface
 
broken cache
 
obvious numerical instability

The speech interface is training.

It is simply not learning a sufficiently useful representation.

Text-QA Positive Control

The frozen B1 text model reaches approximately:

Tune F1 ≈ 0.760

while E1 speech-conditioned performance reaches only:

CF ≈ 0.155
 
XA ≈ 0.098

The text control is not a newly exposure-matched experimental condition, so this should not be treated as a strict causal comparison.

However, diagnostically it strongly suggests:

Qwen QA ability
✓ available
 
speech → Qwen communication
✗ weak

Therefore the central bottleneck is increasingly localized to the interface between speech representations and Qwen.

CF vs XA

At seed17:

CF Clean F1 = 0.155
XA Clean F1 = 0.098

Difference:

CF − XA ≈ +0.057

with a paired bootstrap interval approximately:

[0.019, 0.097]

CF also shows clear same-length speech-content dependence while XA does not.

This suggests that under the current setup:

CF
→ easier to optimize for speech-content dependence
 
XA
→ speech information remains largely unused

However, this is still:

one seed
+
both models failing the grounding gate

Therefore E1 does not establish that CF is scientifically superior to XA.

A meaningful routing comparison still requires both routes to first achieve usable grounding.

Interpretation

E1 further narrows the bottleneck.

Before E1:

Possible explanation:
 
jointly adapting
speech interface + language model
makes grounding difficult

After E1:

capable language model
🔒 frozen
 
speech interface
↓
still fails grounding

Therefore:

Simultaneous language-model adaptation is not the main explanation for the current failure.

The current picture is:

FastConformer speech information
✓ available
 
B1 Qwen QA ability
✓ available
 
CF interface
△ speech-content dependent
  but semantically weak
 
XA interface
✗ little speech-content dependence

The remaining bottleneck is more specifically:

FastConformer representation
↓
projector
↓
language-aligned speech representation
↓
CF / XA
↓
Qwen semantic behavior

What E1 Rules Out

1. Qwen’s QA Capability Is Not the Main Problem

The B1 text model already performs the task substantially better.

Therefore:

Qwen simply cannot answer the questions

is unlikely.

2. Joint Language-Model Adaptation Is Not Necessary for Failure

Even after freezing the language side:

grounding still fails

Therefore the previous failures cannot primarily be blamed on the language model moving while the speech interface is learning.

3. CF Speech Is Not Completely Ignored

CF’s clean-vs-same-length-wrong difference shows that speech content has a measurable effect.

Therefore:

CF route completely ignores speech

is no longer a plausible description.

4. XA Still Does Not Demonstrate Semantic Speech Use

Changing speech content leaves almost all outputs unchanged.

Therefore XA’s remaining behavior cannot yet support any robustness or routing interpretation.

5. More QA CE Alone Is Unlikely to Be the Solution

E1 performs:

1200 ASR
+
1200 QA

updates per route.

Yet both ASR and QA grounding remain poor.

The problem therefore increasingly appears to be the type of supervision, not simply insufficient training duration.

Why B3 Is Still Blocked

B3 aims to study how incoming speech should affect ongoing generation under conditions such as:

interference
 
unreliable speech
 
wrong speaker
 
interruptions

That experiment assumes:

clean speech
↓
both routes understand it

Currently:

CF
→ weak but real speech-content dependence
→ poor semantic accuracy
 
XA
→ almost no speech-content dependence

If B3 were run now:

interference causes little change

could incorrectly be interpreted as:

robust routing

when the actual explanation might simply be:

the route ignores speech

Therefore B3 remains uninterpretable.

Before B3, the project still requires:

  • both routes to pass a common tune grounding gate;

  • replication with seed42;

  • independent listening / reference review;

  • validated generation-interference task;

  • positive controls;

  • explicit B3 authorization.

Conclusion

E1 is a useful failure because it removes another major ambiguity.

The current pipeline can be summarized as:

FastConformer
↓
speech features
✓
 
B1 Qwen
↓
text QA ability
✓
 
CF interface
↓
speech-content dependence
△
 
CF semantic grounding
✗
 
XA speech-content dependence
✗

The central question is now increasingly specific:

How can the projected speech representation be explicitly aligned with a representation that the frozen Qwen language model can reliably interpret?

Ordinary ASR/QA autoregressive CE has not been sufficient.

This motivates stronger direct alignment supervision.

Next Experiment

E2 — Direct-Projector CTC Alignment

E2 branches from each route’s retained:

E1 ASR step1200 checkpoint

Instead of another CE-only QA stage:

speech
↓
projector
↓
CF / XA
↓
Qwen
↓
QA CE

E2 adds an auxiliary direct alignment objective:

                     ┌→ CF / XA → Qwen → QA CE
speech → projector ──┤
                     └→ frozen Qwen vocabulary head
                                      ↓
                                     CTC
                                      ↓
                              spoken question tokens

The important idea is:

Do not only tell the speech interface what final answer Qwen should produce. Explicitly force the projected speech representation to contain lexical information that Qwen’s own vocabulary space can decode.

E1’s CE-only QA branch remains the matched control.

Possible outcomes:

CTC improves
+
grounding improves
↓
weak speech-language alignment was a major bottleneck

or:

CTC improves
+
grounding still fails
↓
lexical information exists,
but routing / Qwen readout cannot use it

or:

CTC itself fails
↓
the speech-to-language interface problem
is even more fundamental

Immediate Goal

Do not enter B3.

The immediate objective remains:

Establish reliable and causal clean-speech grounding for both CF and XA before studying how routing should behave under interference.