2606.13680v1 Jun 11, 2026 cs.CL

검색 기반 강화 학습을 통한 유추 추론 능력 향상

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning

Qi Ma
Qi Ma
Citations: 49
h-index: 2
Vicente Ordonez
Vicente Ordonez
Citations: 57
h-index: 4
Zilin Xiao
Zilin Xiao
Rice University
Citations: 63
h-index: 4
Avinash Atreya
Avinash Atreya
Citations: 40
h-index: 1
Chunliang Chen
Chunliang Chen
Citations: 16
h-index: 2
Hanjie Chen
Hanjie Chen
Rice University
Citations: 2,036
h-index: 16
Xintao Chen
Xintao Chen
Citations: 203
h-index: 9

검색 증강 생성(RAG)은 언어 모델을 외부 지식과 연결하는 표준적인 메커니즘으로 자리 잡았지만, 기존의 어휘 또는 의미 유사성을 기반으로 한 검색 방식은 복잡한 추론 작업에는 적합하지 않습니다. 의미적으로 유사한 문제는 완전히 다른 해결 전략이 필요할 수 있으며, 표면적으로는 다른 문제일지라도 동일한 근본적인 추론 패턴을 공유할 수 있습니다. 본 연구에서는 언어 모델이 유추를 통해 추론하도록 훈련시키는 후처리 프레임워크인 검색 기반 강화 학습 미세 조정(RA-RFT)을 제안합니다. RA-RFT는 '황금 관련성' 증류를 사용하여, 의미 중복보다는 예상되는 추론 효과에 따라 문맥을 순위화하는 검색기를 훈련시킵니다. 그런 다음, 검색된 유사한 예제를 활용하여 강화 학습 미세 조정 방법을 통해 정책 모델을 미세 조정함으로써, 모델은 검증 가능한 결과 보상을 기반으로 추론 과정을 학습합니다. 또한, 검색된 문맥의 다양성을 분석한 결과, 추론에 대한 인식을 갖춘 검색이 상호 보완적인 해결 전략을 제시하며, 개별 문제에 대해 뚜렷한 추론 체계를 제공한다는 것을 확인했습니다. 어려운 수학적 추론 벤치마크에서 RA-RFT는 기존 강화 학습 미세 조정 방법보다 일관되게 우수한 성능을 보였습니다. 예를 들어, RA-RFT는 Qwen3-1.7B 및 Qwen3-4B 모델에서 AIME 2025 평균@32 정확도를 각각 7.1점과 2.8점 향상시켰습니다. 이는 추론에 대한 인식을 갖춘 검색이 보상 설계 또는 교육 커리큘럼의 발전과는 상호 보완적인 개선 요소라는 것을 시사합니다.

Original Abstract

Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic similarity is poorly suited for complex reasoning tasks: a semantically similar problem may demand an entirely different solution strategy, while a superficially different problem may share the same underlying reasoning pattern. We propose Retrieval-Augmented Reinforcement Fine-Tuning (RA-RFT), a post-training framework that teaches language models to reason by analogy. RA-RFT uses gold-relevance distillation to train a retriever that ranks contexts by expected reasoning benefit rather than semantic overlap, and then fine-tunes the policy model via reinforcement fine-tuning methods with retrieved analogous demonstrations, so the model learns to leverage reasoning traces under verifiable outcome rewards. We further analyze the diversity of retrieved contexts and find that reasoning-aware retrieval surfaces complementary solution strategies that provide distinct reasoning scaffolds for individual problems. Across challenging mathematical reasoning benchmarks, RA-RFT consistently outperforms standard reinforcement fine-tuning methods. For example, it improves AIME 2025 average@32 accuracy by 7.1 and 2.8 points over GRPO for Qwen3-1.7B and Qwen3-4B respectively -- suggesting that reasoning-aware retrieval is a complementary axis of improvement and orthogonal to advances in reward design or training curricula.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!