2605.29502v1 May 28, 2026 cs.CL

저자원 대상 언어 생성 모델 학습을 위한 소스 기반 의미론적 강화 학습

Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation

Wentao Zhang
Wentao Zhang
Citations: 20
h-index: 3
Xiaolu Zhang
Xiaolu Zhang
Citations: 1,075
h-index: 5
Ziyin Zhang
Ziyin Zhang
Citations: 404
h-index: 9
Dehan Li
Dehan Li
Citations: 1
h-index: 1
Zhankai Xu
Zhankai Xu
Citations: 42
h-index: 3
Zeli Su
Zeli Su
Citations: 7
h-index: 1
Zewei Pan
Zewei Pan
Citations: 17
h-index: 2
Zhouwu Liu
Zhouwu Liu
Citations: 0
h-index: 0
Di Huang
Di Huang
Citations: 54
h-index: 3
Longfei Zheng
Longfei Zheng
Ant Financial Services Group
Citations: 467
h-index: 11
Jun Zhou
Jun Zhou
Citations: 1,106
h-index: 5

저자원의 대상 언어 생성은 종종 부족한 병렬 데이터에 의해 제한되는 반면, 풍부하지만 표준 지도 학습으로는 활용하기 어려운 고자원 소스 언어 단일 언어 데이터를 보유하고 있습니다. 본 연구에서는 소스 기반 의미론적 강화 학습(SG-SRL)이라는 자원 활용 프레임워크를 제안합니다. SG-SRL은 소스 언어 단일 언어 데이터를 대상 언어 생성에 대한 교차 언어 의미론적 감독 신호로 변환합니다. SG-SRL은 교차 언어 의미론적 보상 모델을 사용하여 소스 언어 데이터에 대해 참조 없이 강화 학습(RL)을 수행하며, 이 보상 모델은 소스 입력과 대상 언어 생성 간의 의미적 관련성을 평가하는 교차 언어 재순위화 모델로 구현됩니다. 이러한 방식은 심각한 장황성 기반의 부정적인 영향을 야기하지만, 작은 병렬 코퍼스를 사용한 경량 복구 단계를 통해 유창성, 간결성 및 작업 형식을 복원하면서 의미론적 이점을 유지합니다. 중국어-태국어 생성 실험 결과, SG-SRL은 콜드 스타트 방식으로 학습된 모델보다 의미론적 근거 및 사실 정보 제공 측면에서 성능이 향상됩니다. 또한 장문 번역 및 티베트어 임베딩 기반 보상에 대한 추가 분석을 통해 SG-SRL의 일반화 동작을 명확히 하고, 실제 저자원 언어 환경에서 인코더 기반 의미론적 보상이 LLM 기반 재순위화 모델을 대체할 수 있음을 보여줍니다.

Original Abstract

Low-resource target-language generation is often limited by scarce parallel data, while high-resource source-language monolingual data is abundant but difficult to use with standard supervised fine-tuning. We propose Source-Grounded Semantic Reinforcement Learning (SG-SRL), a resource-utilization framework that converts source-language monolingual data into cross-lingual semantic supervision for target-language generation. SG-SRL performs reference-free reinforcement learning (RL) on source-language data using a cross-lingual semantic reward model, instantiated by a cross-lingual reranker that scores the semantic relevance between the source input and the target-language generation. While this induces severe verbosity-based reward hacking, a lightweight recovery stage using a small parallel corpus restores fluency, conciseness, and task format while preserving the semantic gains. Experiments on Chinese-to-Thai generation show that SG-SRL improves semantic grounding and factual coverage over cold-start SFT. Additional analyses on long-form transfer and Tibetan embedding-based rewards clarify the generalization behavior of SG-SRL and show that an encoder-based semantic reward can substitute for an LLM-based reranker in a realistic low-resource language setting.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!