DualAnchor: 언어적 선행 지식 보존 및 어휘 정확도 향상을 위한 글로스(Gloss) 없는 수화 번역
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
최근 대규모 언어 모델(LLM)의 발전으로, 수화 번역(SLT), 즉 수화 영상 데이터를 음성 언어 텍스트로 변환하는 작업은 LLM을 텍스트 기반 시스템의 핵심 요소로 점점 더 많이 활용하고 있습니다. 그러나 강력한 언어 모델링 능력을 가지고 있음에도 불구하고, 기존의 LLM 기반 SLT 방법들은 종종 이러한 언어적 선행 지식을 활용하기보다는 저해하며, 이는 어색한 번역으로 이어지는 현상인 '언어적 선행 지식 저하'를 야기합니다. 또한, 기존 방법들은 일반적으로 영상과 텍스트를 문장 수준에서 정렬하는데, 이는 정확한 어휘 정보를 보장하지 못하고 어휘 정확도 격차를 발생시킵니다. 이러한 문제점을 해결하기 위해, 우리는 'DualAnchor'라는 글로스 없는 LLM 기반 SLT 학습 프레임워크를 제안합니다. DualAnchor는 언어적으로 유창하고 시각적으로 충실한 번역을 생성하기 위해 상호 보완적인 두 가지 핵심 요소를 결합합니다. 첫째, 토큰 수준 선행 지식 고정(TPA)은 정해진 LLM의 다음 토큰 분포를 기반으로 다중 모달 디코더를 조정하여 LLM의 언어적 선행 지식을 유지합니다. 둘째, 최적 수송 정렬(OTA)은 시각-텍스트 매칭을 엔트로피 규제된 부분 최적 수송 문제로 정의하고, Sinkhorn 최적화를 통해 시각 토큰과 텍스트 콘텐츠 토큰 간의 부드러운 정렬을 유도하여 어휘 정확도를 향상시킵니다. DualAnchor는 PHOENIX-2014T 및 CSL-Daily 데이터셋 모두에서 뛰어난 성능을 보입니다. 추가 분석 결과, 이러한 성능 향상은 두 가지 요소의 상호 보완적인 효과에 기인합니다. 즉, TPA는 유창성을 개선하고, OTA는 세밀한 수준의 어휘 오류를 줄입니다.
Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM's language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.