Research measuring content and speaker trade-offs when adapting multilingual self-supervised speech models.

The Cost of Speech Adaptation

Continued pre-training can improve speech content tasks while systematically degrading speaker information in discrete-unit models.

Key Takeaways:

Model evaluation icon
Adaptation does not improve every capability.
Problem and risk icon
Discrete-unit models pay the largest speaker cost.
Model behavior icon
Pseudo-labels determine what survives.

Better Domain Adaptation Can Make A Model Worse At Something Else.

Continued pre-training (CPT) gives an existing self-supervised speech model additional unlabeled audio from a target domain. The assumption is intuitive: expose the model to relevant data and its representations should improve.

But representation quality is multidimensional. Factored R&D team studied CPT across HuBERT, WavLM, and OmniASR to understand not only whether adaptation improves multilingual speech performance, but what capabilities may be lost in the process.
‍

Three Ssl Architectures Expose A Hidden Trade-Off.

The study compares two major self-supervised learning paradigms.

HuBERT and WavLM use discrete-unit objectives, predicting offline pseudo-labels at masked positions. OmniASR uses a contrastive wav2vec 2.0-based objective whose targets are generated dynamically.

For the discrete-unit models, the team recomputed k-means pseudo-labels on target-domain audio before resuming masked prediction. OmniASR resumed its native contrastive objective.All three architectures were then evaluated across language identification, automatic speech recognition, and speaker diarization.
‍

Targeted Data Made The Experiment Possible On A Single Gpu.

Instead of attempting to process the full million-hour UPS corpus, the R&D team built a metadata-driven curation pipeline.

Audio shards were scored using voice-activity density and language scarcity, prioritizing speech-rich data while increasing representation from scarcer languages. This produced targeted 100-hour and 500-hour subsets suitable for continued pre-training under constrained compute.

The approach allowed the experiments to run on individual NVIDIA A100 GPUs rather than requiring massive pre-training infrastructure.
‍
‍

Content Improved. Speaker Information Collapsed.

The clearest result emerged in speaker diarization. After continued pre-training:
‍
HuBERT-base: ARI fell from 0.76 → 0.32
WavLM-base+:
ARI fell from 0.59 → 0.31
WavLM-large:
ARI fell from 0.76 → 0.38

OmniASR behaved differently, with ARI increasing slightly from 0.37 → 0.42.

The pattern suggests that discrete-unit continued pre-training shifts representations toward linguistic content at the expense of speaker-discriminative informati

‍

Pseudo-Labels Control What The Model Learns To Preserve.

R&D then changed the pseudo-labeling strategy for HuBERT.

MFCC-based labels degraded content performance but largely preserved speaker information: ARI moved only from 0.76 to 0.72. Embedding-based labels improved content-related metrics but sharply reduced speaker performance, with the primary configuration dropping ARI to 0.32.

The experiments also showed that the transformer layer used to generate those labels can act as another control over the content–speaker balance.

The implication is practical: the target used during continued pre-training helps determine which information the adapted representation prioritizes.
‍

Adaptation Is Specialization, Not A Universal Upgrade.

Continued pre-training should not automatically be interpreted as making a representation universally better. The study found that content-side improvements can vary substantially across runs and data scales. Speaker degradation in discrete-unit models, by contrast, was much more consistent. For teams adapting speech models, that changes the question 

From: “Did the adapted model improve?”

To: “Which capabilities improved—and which ones did we trade away?”

For content-focused applications, CPT can be valuable. But systems that depend on speaker identity or diarization may require additional strategies to preserve that information during adaptation.
‍

Presented at Interspeech 2026.

Forward Deployed Engineering.
In Your Environment.
In Your Time Zone.

Inside your standups, architecture decisions, workflows, and production environments
We build on your stack, for your business
1,000’s of AI & Data engagements across complex production environments
Discuss Your Challenge
Start Building

Continue Reading

Beyond OCR Accuracy
CVPR 2026 Best Short Paper
Structure Beats Random Scale
Language families improve transfer
Speech in the Wild
800K+ hours. 73 languages. 3 tasks.