2607.20090v1 Jul 22, 2026 cs.CL

대규모 언어 모델의 선택적 증거 채택: 오염된 검색 결과로부터의 강화 학습

Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

Yue Li
Yue Li
Citations: 9
h-index: 2
Dongsheng Shi
Dongsheng Shi
Citations: 7
h-index: 2
Yongyi Cui
Yongyi Cui
Citations: 1
h-index: 1
Yanyu Chen
Yanyu Chen
Citations: 48
h-index: 4
Lichang Dai
Lichang Dai
Citations: 0
h-index: 0

검색 기반 대규모 언어 모델은 유용한 정보와 함께 오해를 불러일으키는 내용이나 지시사항과 같은 콘텐츠가 혼합된 맥락에 직면하는 경우가 많습니다. 단순히 모든 정보를 거부하면 유효한 증거를 버리는 반면, 비판적으로 검토하지 않고 채택하면 부정확하거나 위험한 답변을 생성할 수 있습니다. 따라서 실제 검색 환경에서 안정적인 성능을 위해서는 관련 정보는 선택적으로 채택하고, 기만적이거나 해로운 콘텐츠는 거부하는 능력이 매우 중요합니다. 본 연구에서는 선택적 증거 채택을 위한 제어된 벤치마크 및 학습 데이터셋인 SelectBench를 소개하고, DAPO(Deterministic Action Programming Optimization) 방식을 사용하여 Qwen3.5-4B 모델을 결정적인 규칙 기반 보상 또는 동결된 의미론적 판단자를 통해 추가적으로 학습했습니다. 수정된 325개 예시로 구성된 SelectBench-v2 테스트 데이터셋에서, 엄격한 성공률은 원래 체크포인트의 22.46%에서 DAPO-Rule을 사용했을 때 25.54%, DAPO-DeepSeek을 사용했을 때는 26.46%로 향상되었습니다. 두 가지 학습된 모델 모두 금지된 콘텐츠 채택 비율을 줄이고, 더 짧고 집중적인 답변을 생성하지만, 프롬프트 주입 공격에 대한 저항성은 개선되지 않았습니다. 이러한 개선 효과는 미미하며, Holm 보정을 통해 유의미하지 않은 것으로 나타났습니다. 따라서 더욱 강력한 보상 설계 또는 추가적인 학습 반복이 필요할 수 있습니다. DAPO-DeepSeek은 MMLU 및 깨끗한 HotpotQA 데이터셋에서 현저한 성능 저하를 보이지 않아, 추가 학습 과정이 모델의 일반적인 능력을 유지한다는 것을 시사합니다. 이러한 결과는 선택적 증거 활용 능력에 대한 긍정적인 방향성을 보여주며, 프롬프트 주입 방지 및 통계적 견고성 확보가 향후 연구에서 중요한 과제임을 강조합니다.

Original Abstract

Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect or unsafe answers. The ability to selectively adopt relevant information while rejecting deceptive or harmful content is therefore critical for reliable deployment in real-world retrieval settings. We introduce SelectBench, a controlled benchmark and training set for selective evidence adoption, and post-train Qwen3.5-4B directly with DAPO using either deterministic rule rewards or a frozen semantic judge. On the corrected 325-example SelectBench-v2 test set, strict success rises from 22.46% for the original checkpoint to 25.54% with DAPO-Rule and 26.46% with DAPO-DeepSeek. Both trained policies reduce forbidden-content adoption and produce shorter, more focused responses, yet prompt-injection following does not improve. The paired gains are modest and fail to survive Holm correction, suggesting that stronger reward shaping or additional training iterations may be needed for more robust gains. DAPO-DeepSeek exhibits no material degradation on MMLU or clean HotpotQA, indicating that the post-training procedure preserves general capabilities. These results demonstrate a directional improvement in selective evidence use, while identifying injection resistance and statistical robustness as important remaining challenges for future work.

1 Citations
0 Influential
2 Altmetric
11.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!