AI & Computational Science

Auditing Cross-Lingual Fairness in Language Model Watermarking

How the science connects

Natural language p…Digital watermarking

AI Insight

This study evaluates watermarking techniques for large language model outputs across multiple languages and reveals significant fairness issues that are hidden when testing only on English. The researchers tested six watermarking schemes across eleven languages from different language families and found that detection accuracy and text quality degradation vary systematically based on structural linguistic properties rather than individual language characteristics. Their analysis shows that current watermarking methods perform inconsistently across languages, with disparities aligning more with typological language families than with specific languages.


As language models are deployed globally, watermarking systems that work well in English may fail or produce lower-quality text in other languages, creating inequitable access and utility. This work provides a framework for evaluating cross-lingual fairness in watermarking, which is essential for ensuring that content authentication and AI detection tools work equitably across linguistic communities.


Understand the Science

Natural language processing 57 articles Explore Concept → Digital watermarking Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme’s detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.

Source: Auditing Cross-Lingual Fairness in Language Model Watermarking