2607.25268v1 Jul 28, 2026 cs.IR

순위 결정에 대한 구조 인식 상대 정책 최적화

Structure-aware Relative Policy Optimization for Ranking

Yiqun Liu
Yiqun Liu
Citations: 1,716
h-index: 22
Qingyao Ai
Qingyao Ai
Citations: 1,764
h-index: 22
Weihang Su
Weihang Su
Citations: 797
h-index: 18
Min Zhang
Min Zhang
Citations: 104
h-index: 6
Yiteng Tu
Yiteng Tu
Citations: 59
h-index: 3
Zitao Su
Zitao Su
Citations: 0
h-index: 0

순위 결정은 현대 정보 접근 시스템의 핵심 구성 요소입니다. 강화 학습(RL)은 전체 순위 목록에 대해 정의된 거칠고 추상적인 피드백과 시스템 수준 목표를 직접적으로 최적화하는 유연한 프레임워크를 제공합니다. 그러나 기존의 RL 기반 순위 결정 방법은 일반적으로 각 샘플링된 순열을 하나의 독립적인 출력으로 취급하고, 주로 스칼라 보상을 통해 평가하며, 서로 다른 순위 목록 간의 구조적 관계를 간과합니다. 결과적으로, 유사한 보상을 갖지만 순열 패턴이 크게 다른 순열들이 동일한 최적화 신호를 받을 수 있으며, 이는 부정확한 기여도 할당 및 과도하게 공격적인 정책 업데이트로 이어질 수 있습니다. 이러한 한계를 해결하기 위해, 우리는 목록 기반 순위 결정에 대한 구조 인식 상대 정책 최적화(SRPO) 프레임워크를 제안합니다. SRPO는 샘플링된 순열 간의 차이를 상위 항목 가중 켄달-타우 거리(Kendall's tau distance)를 사용하여 측정하고, 해당 거리에 따라 쌍별 보상 차이를 정규화합니다. 이를 통해 순위 변화 단위당 보상 향상을 정량화하여 효율적인 지역 최적화를 강조하며, 특히 상위 순위에 관련된 개선 사항에 중점을 둡니다. 두 가지 순위 결정 시나리오에서의 실험 결과는 순열 수준의 차이를 명시적으로 모델링하면 목록 기반 순위 결정을 더욱 효과적이고 안정적으로 만들 수 있다는 것을 보여주며, 특히 제한된 피드백 및 복잡한 목록 수준 최적화 환경에서 우수한 성능을 보입니다.

Original Abstract

Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for directly optimizing coarse-grained feedback and system-level objectives defined over the complete ranking list. However, existing RL-based ranking methods typically treat each sampled permutation as an atomic output and evaluate it primarily through a scalar reward, overlooking the structural relationships among different ranking lists. Consequently, permutations with similar rewards but substantially different permutation patterns may receive comparable optimization signals, potentially leading to inaccurate credit assignment and overly aggressive policy updates. To address this limitation, we propose SRPO, a \textbf{S}tructure-aware \textbf{R}elative \textbf{P}olicy \textbf{O}ptimization framework for listwise ranking. SRPO measures the discrepancy between sampled permutations using a top-weighted Kendall-tau distance and normalizes their pairwise reward differences by the corresponding distances. It quantifies the reward improvement per unit of ranking change, thereby emphasizing efficient local refinements, particularly those involving top-ranked positions. Experimental results across two ranking scenarios demonstrate that explicitly modeling permutation-level differences improves the effectiveness and stability of listwise ranking, with particularly favorable performance in limited-feedback and complex list-level optimization settings.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!