2608.06377v1 Aug 06, 2026 cs.CL

선택적 문맥 선호도 최적화를 통한 신뢰 시점 학습

Learning When to Trust via Selective Context Preference Optimization

Junhao Liu
Junhao Liu
Citations: 13
h-index: 2
Yingshuo Wang
Yingshuo Wang
Citations: 6
h-index: 1
Wei Chow
Wei Chow
Citations: 404
h-index: 8
Wei Gao
Wei Gao
Citations: 20
h-index: 2
Lingdong Kong
Lingdong Kong
Citations: 151
h-index: 7
Xianshui Sun
Xianshui Sun
Citations: 0
h-index: 0
Qingan Wu
Qingan Wu
Citations: 3
h-index: 1

언어 모델은 점점 더 외부 신호에 의존하여 답변을 생성하며, 단 하나의 잘못된 신호가 올바른 답변을 틀린 답변으로 만들 수 있습니다. 이러한 문제에 대한 명백한 해결책은 모델이 그러한 신호에 저항하도록 훈련하는 것입니다. 하지만 이는 또 다른 문제점을 숨깁니다. 모든 문맥을 무시하는 모델은 견고해 보이지만, 실제로 신뢰할 가치가 있는 문맥이 있을 때 유용하지 않습니다. 우리는 이 문제를 '선택적 신뢰'라는 관점에서 재정의하고, MIST라는 인간 주석 데이터셋을 소개합니다. MIST는 각 추론 항목을 네 가지 조건(깨끗한 상태, 오해를 불러일으키는 상태, 올바른 문맥 상태, 관련 없는 문맥 상태)으로 구성하며, SC2W라는 지표를 통해 오해를 불러일으키는 신호가 깨끗하고 올바른 답변을 틀린 답변으로 바꾸는 빈도를 측정합니다. 종합적인 벤치마크 연구 결과, 이러한 취약성은 보편적으로 나타나는 것을 확인했습니다. 우리는 또한 SCOPE라는 방법을 제안하는데, 이는 깨끗하고 올바른/오해를 불러일으키는-틀린 사례의 실패를 분석하고, 표준 Direct Preference Optimization (DPO) 목적 함수를 네 가지 조건에 균등하게 분배된 쌍으로 구성된 선호도 쌍을 사용하여 최적화합니다. 우리의 접근 방식은 인기 있는 오픈 소스 모델에서 SC2W 값을 크게 줄이는 동시에 추가된 문맥이 깨끗하거나 올바르거나 관련 없는 경우 정확도를 유지합니다. 본 연구를 통해 우리는 모델의 성능을 저항 능력만으로 평가하는 것이 아니라, '선택적 신뢰' 능력을 기준으로 평가해야 한다고 주장합니다.

Original Abstract

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open-sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!