Research on robust multilingual speech representations trained on large-scale, real-world web audio.

Speech in the Wild

A new benchmark tests multilingual speech models against 800K+ hours of heterogeneous, real-world web audio.

Key Takeaways:

800K+ hours bring speech AI into the wild.
Benchmark evaluation icon
One benchmark tests three dimensions of speech.
No single model wins everywhere.

Real-World Speech Breaks the Assumptions of Curated Data.

Self-supervised learning has transformed speech AI by allowing models to learn powerful representations from audio without transcripts. But much of that progress has been built and evaluated using comparatively curated speech corpora.

Real-world audio is different.

It contains spontaneous conversation, background music, environmental noise, inconsistent recording conditions, and long-tailed linguistic variation. The Unsupervised Speech in the Wild (UPS) 2026 Challenge was created to test whether speech representations can remain useful under those conditions.

800,000+ Hours Put Speech Models Under Real-World Pressure.

The challenge restricts training to the Unsupervised People’s Speech dataset, a massive collection of publicly licensed web audio.

The dataset contains more than 800,000 hours of audio, with approximately 522,000 hours of detected speech. Analysis identified 89 languages, creating a training environment far more acoustically and linguistically heterogeneous than conventional curated speech datasets.

The challenge asks a fundamental question:

Can self-supervised models learn representations that generalize when their training data actually looks like the messy, multilingual audio found in the real world?

Three Tasks Test What the Representations Actually Learn.

Rather than optimizing for one downstream benchmark, UPS evaluates the same frozen speech representation across three complementary tasks.

Language Identification
Models identify speech across 73 languages, measuring whether representations encode language-discriminative information.

Few-Shot Automatic Speech Recognition
Models perform character-level transcription across the same 73 languages using limited labeled data, testing how well linguistic and phonetic information transfers.

Speaker Clustering
Models distinguish 398 speakers across 70 languages, measuring whether their representations retain speaker-level information independent of linguistic content.

A unified encoder interface and standardized downstream probes keep the comparison focused on the quality of the learned representation rather than task-specific engineering.

80 Submissions Revealed There Is No Universal Winner.

The challenge produced 80 successfully scored submissions representing 79 unique model IDs from 14 teams.

The results exposed substantial differences between models—and a critical finding: no single model dominated every task.

Whisper and Nx models produced the strongest language-identification results. WavLM-large variants achieved the lowest character error rates for ASR. A Qwen3 encoder baseline delivered the strongest speaker-clustering result.

That separation matters.

It suggests language, transcription, and speaker information represent partially distinct capabilities inside learned speech representations. Optimizing one does not guarantee strength across the others.

Language Itself Exposes Where Evaluation Breaks Down.

Performance also varied substantially by language.

For ASR, languages using non-Latin or more complex writing systems proved particularly difficult under character-level evaluation. Japanese produced the highest mean character error rate, while Korean, Amharic, Khmer, and Thai were also among the most difficult.

Speaker clustering revealed another challenge: dataset composition matters. The number and balance of speakers within individual languages could materially affect clustering performance.

These results show that multilingual benchmarks cannot be understood through aggregate scores alone.

Robust Speech AI Needs Robust Benchmarks.

The UPS Challenge moves speech representation evaluation closer to the conditions models encounter outside controlled datasets.

Its results demonstrate meaningful progress in learning from heterogeneous web audio while exposing limitations that single-task or heavily curated benchmarks can hide.

Building speech systems that work across languages, speakers, and environments requires evaluating all three—not optimizing one in isolation.

Presented at Interspeech 2026.

Forward Deployed Engineering.
In Your Environment.
In Your Time Zone.

Inside your standups, architecture decisions, workflows, and production environments
We build on your stack, for your business
1,000’s of AI & Data engagements across complex production environments
Discuss Your Challenge
Start Building

Continue Reading

Architecting Trust in AI Agents
91% completion, reasoning still fails
Medical LLMs: Real-World Risks
1,298-person study reveals reliability gaps
Klingon Effect In Multilingual AI
Rare-language data boosts robustness