2605.26554v1 May 26, 2026 cs.LG

지연 피드백을 사용하는 선형 및 신경망 다중 액터 강화 학습

Linear and Neural Dueling Bandits with Delayed Feedback

Mingze Kong
Mingze Kong
Citations: 17
h-index: 3
Zhi Hong
Zhi Hong
Citations: 3,281
h-index: 3
Zhiyong Wang
Zhiyong Wang
Citations: 28
h-index: 3
Zhongxiang Dai
Zhongxiang Dai
Citations: 65
h-index: 5
Jieming Mao
Jieming Mao
Citations: 32
h-index: 4
Xiangyi Wang
Xiangyi Wang
Citations: 6
h-index: 1
Pingchen Lu
Pingchen Lu
Citations: 0
h-index: 0

문맥 기반 다중 액터 강화 학습은 추천 시스템 및 대규모 언어 모델 정렬과 같은 영역에서 중요한 역할을 합니다. 그러나 기존 알고리즘은 종종 현실 세계 시나리오, 특히 프롬프트 최적화와 같이 지연 피드백이 발생하는 경우에도 적용될 수 있는 이상적인 즉각적인 피드백이라는 가정에 의존합니다. 이러한 환경은 독특한 이론적 과제를 제시하며, 선형 강화 학습과 달리 다중 액터 강화 학습 추정기는 해석적인 해를 갖지 않아, 표준 가중치 기법의 단순한 적용 방식이 편향을 초래할 수 있습니다. 이를 해결하기 위해, 우리는 확률적 지연 피드백을 포함하는 문맥 기반 다중 액터 강화 학습 문제를 공식화하고, 두 가지 새로운 알고리즘인 지연 피드백을 사용하는 선형 다중 액터 강화 학습 (LDB-DF) 및 신경망 다중 액터 강화 학습 (NDB-DF)을 제안합니다. 우리의 접근 방식의 핵심은 손실 함수에 역 확률 가중치 (IPW) 메커니즘을 직접 통합하여 지연되거나 누락된 피드백으로 인한 편향을 수정하는 새로운 추정기를 사용하는 것입니다. 우리는 포괄적인 이론적 분석을 제공하며, 선형 설정에서는 O(d*sqrt(T))의 후회 경계를, 신경망 설정에서는 부분 선형 보장을 제시합니다. 시뮬레이션 및 실제 데이터 세트에서 수행한 광범위한 실험은 제안된 알고리즘의 효과를 입증했습니다.

Original Abstract

Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large language model alignment. However, standard algorithms rely on the idealized assumption of immediate feedback, a condition frequently violated in real-world scenarios such as prompt optimization. This setting introduces a unique theoretical challenge: unlike linear bandits, dueling bandit estimators lack closed-form solutions, rendering naive adaptations of standard weighting techniques biased. To address this, we formalize the problem of Contextual Dueling Bandits with Stochastic Delayed Feedback and propose two novel algorithms: Linear (LDB-DF) and Neural (NDB-DF) Dueling Bandits with Delayed Feedback. Central to our approach is a novel estimator that integrates an Inverse Probability Weighting (IPW) mechanism directly into the loss function, ensuring unbiased correction for delayed or missing feedback. We provide comprehensive theoretical analysis, establishing an O(d*sqrt(T)) regret bound for the linear setting and sub-linear guarantees for the neural setting. Extensive experiments on both simulated and real-world datasets demonstrate the effectiveness of our propose.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!