ALSACE: Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation

← All research

Language Understanding

NAACL 2024 · PDF · Paper · Code

Let languages within the same model teach one another. Multilingual pretrained language models often answer the same question differently across languages. ALSACE addresses this knowledge and performance disparity without requiring additional labeled multilingual data.

ALSACE example showing predictions for the same cultural question in English, Chinese, Hindi, Persian and Swahili.
Figure 1: A GeoMLAMA example illustrates how sharing knowledge across languages can improve their answers to the same question. Click to enlarge.

It first uses cross-language majority voting to form pseudo-labels and selects reliable teacher languages using their confidence on those labels. Cross-lingual self-distillation then aligns prediction distributions between selected teachers and other languages. The teachers adapt to the task, so useful knowledge can flow from both high-resource and low-resource languages.

The paper evaluates XNLI, PAWS-X and XCOPA, and probes cultural knowledge with GeoMLAMA, using XLM-R and mT5 models. In these experiments, ALSACE reduces cross-lingual transfer gaps while improving performance across many languages, including limited-resource settings. Its gains come from sharing knowledge already present in the model using unlabeled multilingual inputs.