AI & Computational Science

Phonological Interference in Multilingual Speech Models

How the science connects

Speech recognitionMultilingualismPhonology

AI Insight

This study identifies a systematic failure in multilingual speech models called "phonological interference," where models incorrectly assume input belongs to a single language and impose that language's sound rules throughout, even when processing code-switched speech or unfamiliar languages. Researchers found that phone recognizers and text-to-speech models lose 32% to 79% of phonemes unique to one language when processing code-switched input, while phonemes shared between languages remain largely intact. The team developed an inference-time correction method called windowed language estimation (WLE) that reduces interference by 34% to 69% across tested models without degrading monolingual performance.


These findings address critical limitations in multilingual speech technology that affect billions of multilingual speakers who naturally code-switch between languages. The proposed WLE method could improve speech recognition and synthesis for multilingual contexts and low-resource languages without requiring model retraining, making speech technology more accessible and accurate for diverse linguistic communities.


Understand the Science

Speech recognition Concept coming soon Multilingualism Concept coming soon Phonology Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models’ language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model’s language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.

Source: Phonological Interference in Multilingual Speech Models