2605.25850v1 May 25, 2026 cs.CL

TIAR: 경로 정보를 활용한 장점 재가중치를 통한 LLM 회피 학습

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

Muyu Pan
Muyu Pan
Citations: 3
h-index: 1
Shu Zhao
Shu Zhao
Citations: 20
h-index: 2
Nan Zhang
Nan Zhang
Citations: 471
h-index: 6
Philip Shin
Philip Shin
Citations: 45
h-index: 2
Varun Parekh
Varun Parekh
Citations: 2
h-index: 1
Rui Zhang
Rui Zhang
Citations: 54
h-index: 3
N. Vijaykrishnan
N. Vijaykrishnan
Citations: 19,562
h-index: 71

본 논문은 대규모 언어 모델(LLM)의 회피 학습을 연구하며, 특히 LLM의 진실성을 유도하는 3진 시스템 보상을 사용합니다. 본 논문에서는 이 아이디어를 확장하여 그룹 상대 정책 최적화(GRPO) 훈련 과정에서 경로 정보를 활용한 장점 재가중치 방식을 도입하여 회피 보상을 동적으로 조정합니다. 본 연구의 목표는 진실성 향상보다는 회피 학습에 초점을 맞추어 환각 현상 감소를 위한 탐색을 수행하는 것입니다. 본 논문의 차별성은 방법론적 혁신, 장점 재가중치 방식 및 벤치마크 선정에 있습니다. GRPO의 다수의 경로를 자연스러운 회피 신호로 활용하여, 보상 신호를 통해 지식 경계를 탐색하고 일관성을 장려합니다. 본 연구는 경로가 정책의 신뢰도를 나타내는 지표로 사용될 수 있음을 보여주고, 이를 바탕으로 동적으로 회피 관련 이점을 계산합니다. AbstentionBench를 평가 벤치마크로 사용하여, 회피 학습 분야에 기여하고자 합니다. 벤치마크에 포함된 모든 데이터셋을 본 방법과 다양한 기준 모델에 대해 테스트했습니다. 실험 결과, TIAR는 6개의 평가 항목 중 5개에서 최고 수준의 회피 F1 점수를 달성했으며, 31개의 벤치마크 데이터셋 중 17개에서 기준인 3진 시스템보다 더 우수한 성능을 보였으며, 동시에 기준 모델의 정확도를 그대로 유지했습니다.

Original Abstract

This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large language models. This paper extends that idea by moving from a ternary reward to a Trajectory-Informed advantage reweighting, dynamically re-weights the abstention reward during Group Relative Policy Optimization (GRPO) training. The objective of this work focuses on abstention learning instead of improving truthfulness, serving as an exploration into hallucination reduction. The novelty of this paper lies in methodological innovation, advantage re-weighting, and benchmark selection. Leveraging GRPO's multiple trajectories as a natural abstention signal, this method uses a reward signal to explore knowledge boundaries and encourage consistency. By demonstrating that trajectories can be used as a confidence indicator of the policy relative to the query, they are then used to dynamically calculate the abstention advantage. AbstentionBench is used as the evaluation benchmark, as this work aims to contribute to the field of abstention learning. All datasets on the benchmark were tested against this method and various baselines. Empirical results demonstrate that TIAR achieves state-of-the-art abstention F1 scores across five of six evaluation categories, outperforming the static ternary baseline on 17 of 31 benchmark datasets while fully preserving baseline accuracy.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!