2605.05745v1 May 07, 2026 cs.AI

하이브리드 피드백을 이용한 일반화 선형 밴딧에서의 최적 팔 식별

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback

Xuchuang Wang
Xuchuang Wang
Citations: 158
h-index: 8
Qirun Zeng
Qirun Zeng
Citations: 3
h-index: 1
Xutong Liu
Xutong Liu
University of Washington
Citations: 390
h-index: 14
Fang-yuan Kong
Fang-yuan Kong
Citations: 222
h-index: 9
Jinhang Zuo
Jinhang Zuo
Citations: 352
h-index: 11
Jiayi Shen
Jiayi Shen
Citations: 42
h-index: 2

본 연구는 일반화 선형 밴딧 환경에서, 각 라운드마다 학습자가 (i) 단일 팔로부터 절대 보상 피드백 또는 (ii) 팔 쌍으로부터 상대적(대결) 피드백을 받을 수 있는 하이브리드 피드백 모델 하에서, 고정된 신뢰 수준을 갖는 최적 팔 식별 문제를 다룬다. 제시된 접근 방식은 일반화 선형 모델에 의해 관리되는 이질적인 일반화 선형 관측치를 통합하는 likelihood-ratio 기반 신뢰 시퀀스를 도입하며, 자체 조화(self-concordance) 가정 하에서 명시적인 타원형 신뢰 집합을 제공한다. 이 신뢰 집합을 기반으로, 학습자는 팔과 팔 쌍의 공동 행동 공간에서 minimax-최적의 설계를 추적하여 쿼리를 적응적으로 할당하는 하이브리드 Track-and-Stop 알고리즘을 제안한다. 본 연구는 δ-정확성을 입증하고 중지 시간의 고확률 상한을 제시하며, 피드백 방식에 따른 이질적인 획득 비용을 고려하는 비용 민감 설정으로 프레임워크를 확장한다. 실험 결과, 제안된 알고리즘이 기존 방법보다 샘플 효율성을 크게 향상시키는 것으로 나타났다.

Original Abstract

We study fixed-confidence best arm identification in generalized linear bandits under a hybrid feedback model: at each round, the learner may query either (i) absolute reward feedback from a single arm or (ii) relative (dueling) feedback from an arm pair, both governed by generalized linear models. We introduce a likelihood-ratio--based confidence sequence that unifies heterogeneous generalized linear observations and yields an explicit ellipsoidal confidence set under a self-concordance assumption. Building on this confidence set, we propose a hybrid Track-and-Stop algorithm that adaptively allocates queries by tracking a minimax-optimal design over a joint action space of arms and pairs. We establish $δ$-correctness and provide high-probability upper bounds on the stopping time. We further extend the framework to a cost-aware setting that accounts for heterogeneous acquisition costs across feedback modalities. Empirical experiments demonstrate that the proposed algorithms significantly improve sample efficiency over baseline methods.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!