다국어 병렬 텍스트 정렬을 통한 다국어 임베딩 성능 향상
Enhancing Multilingual Embeddings via Multi-Way Parallel Text Alignment
다국어 사전 학습은 일반적으로 명시적인 정렬 신호를 포함하지 않아, 표현 공간에서의 최적의 교차 언어 정렬을 달성하지 못하는 경우가 많습니다. 본 연구에서는 다양한 언어 집합에 대한 다국어 병렬 코퍼스를 사용하여 표준 사전 학습 모델을 교차 언어 정렬에 맞게 훈련하면, 자연어 이해(NLU) 작업에 대한 다국어 및 교차 언어 표현 능력을 크게 향상시킬 수 있음을 보여줍니다. 저희는 상용 신경망 기계 번역(NMT) 모델을 사용하여 영어 텍스트를 여섯 개의 대상 언어로 번역한 다국어 병렬 데이터셋을 구축하고, 대비 학습을 통해 강력한 교차 언어 정렬을 달성했습니다. 이는 XLM-Roberta 및 다국어 BERT base 모델에 대한 MTEB 벤치마크 평가에서, 기존 언어와 새로운 언어 모두에서 상당한 성능 향상을 가져왔습니다. 대비 학습을 위한 다국어 병렬 코퍼스 사용은 영어 중심(En-X)의 이중 언어 병렬 데이터에 비해 비텍스트 마이닝(21.3%), 의미 유사성(5.3%), 분류(28.4%) 작업에서 상당한 성능 향상을 가져왔습니다. 또한, 다국어 병렬성을 갖춘 작은 데이터셋으로 mE5 모델을 미세 조정하면, 병렬성이 없는 경우보다 비텍스트 마이닝 성능이 크게 향상되었으며, 이는 고품질 문장 임베딩을 위해 사전 학습된 모델에서도 다국어 교차 언어 감독 학습의 중요성을 강조합니다.
Multilingual pretraining typically lacks explicit alignment signals, leading to suboptimal cross-lingual alignment in the representation space. In this work, we show that training standard pretrained models for cross-lingual alignment with a multi-way parallel corpus in a diverse pool of languages can substantially improve multilingual and cross-lingual representations for NLU tasks. We construct a multi-way parallel dataset using translations of English text from an off-the-shelf NMT model for a pool of six target languages and achieve strong cross-lingual alignment through contrastive learning. This leads to substantial performance gains across both seen and unseen languages for multiple tasks from the MTEB benchmark evaluated for XLM-Roberta and multilingual BERT base models. Using a multi-way parallel corpus for contrastive training yields substantial gains on bitext mining (21.3%), semantic similarity (5.3%), and classification (28.4%) compared to English-centric (En-X) bilingually parallel data, where X is sampled from a pool of multiple target languages. Furthermore, finetuning mE5 model on a small dataset with multi-way parallelism significantly improves bitext mining compared to one without, underscoring the importance of multi-way cross-lingual supervision even for models already pretrained for high-quality sentence embeddings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.