2608.04586v1 Aug 05, 2026 cs.CL

다대다 음성-텍스트 번역에서 리소스 기반 혼합 음성 인코더를 통한 다국어 문제 해결

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Youcheng Pan
Youcheng Pan
Citations: 212
h-index: 8
Yexing Du
Yexing Du
Citations: 97
h-index: 4
Kaiyuan Liu
Kaiyuan Liu
Citations: 44
h-index: 4
Bo Yang
Bo Yang
Citations: 111
h-index: 4
Ming Liu
Ming Liu
Citations: 35
h-index: 3
Chengpeng Fu
Chengpeng Fu
Citations: 24
h-index: 2
Yu Wang
Yu Wang
Citations: 0
h-index: 0

멀티모달 대규모 언어 모델(MLLM)은 음성-텍스트 번역(S2TT) 분야에서 상당한 성공을 거두었습니다. 그러나 다국어 음성 입력을 처리할 때, 모든 언어에 공유되는 단일 음성 인코더는 '다국어의 저주' 현상을 겪습니다. 즉, 서로 다른 수준의 리소스를 가진 언어들이 제한된 표현 능력을 놓고 경쟁하여, 풍부한 자원을 가진 언어에서는 높은 성능을 보이지만, 자원이 부족한 음성에 대해서는 성능이 크게 저하됩니다. 이러한 문제를 해결하고 다국어 일관성을 향상시키기 위해, 본 연구에서는 리소스 기반 혼합 음성 인코더(MoSE)를 중심으로 구축된 새로운 프레임워크인 MSRT를 제안합니다. MoSE는 명시적인 언어 라우터를 사용하여 각 발화를 적절한 전문 인코더에 할당합니다. 고정된 전문 인코더는 풍부한 자원을 가진 언어의 능력을 유지하는 반면, 학습 가능한 전문 인코더는 중간 및 저자원 언어에 적응하고 특화됩니다. 또한, 데이터 의존성을 크게 줄이는 5단계 커리큘럼 학습 전략을 도입하여, 각 언어별로 단 10시간의 쌍방향 S2TT 데이터를 사용하여 효과적인 정렬을 수행합니다. 우리는 총 45개 언어를 대상으로 광범위한 실험을 진행하여 모든 45 x 44번의 번역 방향에 대해 체계적으로 평가했습니다. 40억 개의 파라미터를 가진 우리의 모델은 최고 수준의 성능을 달성했으며, 훨씬 더 큰 기본 모델보다 우수한 결과를 보였습니다. 실증적 분석 결과, MoSE는 고자원, 중자원 및 저자원 언어 모두에서 동시에 성능을 향상시키며, 특히 저자원 음성에 대한 개선 효과가 가장 큽니다. 이를 통해 다국어의 저주를 극복하고, 동시에 고자원의 성능 저하 없이 멀티모달 S2TT 연구를 지원하기 위해, 본 연구의 코드 및 모델을 공개합니다.

Original Abstract

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistency, we propose MSRT, a novel framework built around a resource-aware Mixture of Speech Encoders (MoSE). MoSE uses an explicit language router to assign each utterance to an appropriate expert encoder. A frozen expert preserves high-resource language capabilities, while a trainable expert adapts to and specializes in medium- and low-resource languages. We further introduce a five-stage curriculum learning strategy that substantially reduces data dependence, requiring only 10 hours of paired S2TT data per language for effective alignment. We conduct extensive experiments on 45 languages, systematically evaluating all $45 \times 44$ translation directions. Our 4B-parameter model achieves state-of-the-art performance, outperforming substantially larger baselines. Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance. To support future multilingual S2TT research, we release our code and models.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!