2605.28642v1 May 27, 2026 cs.AI

대역폭 효율적이고 개인 정보 보호를 위한 엣지-클라우드 다대다 음성 번역

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

Bing Qin
Bing Qin
Citations: 682
h-index: 12
Ming Liu
Ming Liu
Citations: 627
h-index: 12
Youcheng Pan
Youcheng Pan
Citations: 212
h-index: 8
Yexing Du
Yexing Du
Citations: 97
h-index: 4
Kaiyuan Liu
Kaiyuan Liu
Citations: 44
h-index: 4
Bo Yang
Bo Yang
Citations: 111
h-index: 4
Yang Xiang
Yang Xiang
Citations: 40
h-index: 3

멀티모달 대규모 언어 모델(MLLM)은 음성-텍스트 번역(S2TT)에서 상당한 잠재력을 보여주었습니다. 그러나 기존의 배포 방식은 심각한 문제에 직면합니다. 순수 온디바이스 모델은 리소스 제약으로 인해 어려움을 겪고, 중앙 집중식 클라우드 시스템은 원시 음성 데이터를 전송함으로써 심각한 개인 정보 위험과 대역폭 병목 현상을 초래합니다. 또한, 대부분의 모델은 영어 중심적인 편향을 나타내어 다대다 번역 확장에 제약을 가합니다. 본 논문에서는 개인 정보를 보호하고 대역폭 효율성을 높인 협업 엣지-클라우드 MLLM 프레임워크인 Edge-cloud Speech Recognition and Translation (ESRT)을 제안합니다. 구체적으로, 우리는 장치에 경량 음성 인코더와 어댑터를 유지하고, 압축된 중간 특징 데이터만 클라우드로 전송하는 엣지-클라우드 분산 추론 아키텍처를 설계했습니다. 이를 통해 음성 정보 유출을 근본적으로 방지하고 대역폭 요구 사항을 최대 10배까지 줄일 수 있습니다. 영어 중심적인 병목 현상을 극복하기 위해, 데이터 균형을 고려한 다중 작업 가중 학습 전략을 도입하여 강력한 교차 언어 일관성을 보장합니다. FLEURS 데이터 세트에 대한 광범위한 실험 결과, ESRT-4B 및 ESRT-12B 모델이 45개 언어(45 x 44 방향)에 걸쳐 최첨단 다대다 S2TT 성능을 달성하는 것으로 나타났습니다. 코드와 모델은 재현 가능한 개인 정보 보호 MLLM S2TT 연구를 촉진하기 위해 공개됩니다 (https://github.com/yxduir/esrt).

Original Abstract

Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur severe privacy risks and bandwidth bottlenecks by transmitting raw voice data. Furthermore, most models exhibit English-centric biases, restricting many-to-many translation scaling. In this paper, we propose Edge-cloud Speech Recognition and Translation (ESRT), a privacy-preserving and bandwidth-efficient collaborative edge-cloud MLLM framework. Specifically, we design an edge-cloud split inference architecture that retains a lightweight speech encoder and adapter on the device, transmitting only highly compressed intermediate features to the cloud. This fundamentally prevents voiceprint leakage and reduces bandwidth requirements by up to 10$\times$. To overcome English-centric bottlenecks, we introduce a multi-task weighted curriculum learning strategy with data balancing to ensure robust cross-lingual consistency. Extensive experiments on the FLEURS dataset demonstrate that our models, ESRT-4B and ESRT-12B, achieve state-of-the-art many-to-many S2TT performance across 45 languages ($45 \times 44$ directions). Code and models are released to facilitate reproducible, privacy-aware MLLM S2TT research. The code and models are released at https://github.com/yxduir/esrt.

0 Citations
0 Influential
29.4657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!