2601.19030v1 Jan 26, 2026 cs.LG

선형 오프라인 평가에서의 보장(Coverage)에 대한 통합적 관점

A Unifying View of Coverage in Linear Off-Policy Evaluation

P. Amortila
P. Amortila
Citations: 297
h-index: 8
Audrey Huang
Audrey Huang
Citations: 17
h-index: 3
Akshay Krishnamurthy
Akshay Krishnamurthy
Citations: 466
h-index: 10
Nan Jiang
Nan Jiang
Citations: 94
h-index: 4

오프라인 평가(OPE)는 강화 학습(RL)의 핵심적인 과제입니다. 선형 OPE의 전형적인 설정에서, 유한 표본 보장은 종종 다음과 같은 형태로 나타납니다: $$ extrm{평가 오차} extless{} extrm{poly}(C^π, d, 1/n, ext{log}(1/δ)), $$ 여기서 $d$는 특징의 차원이며, $C^π$는 방문된 특징들이 데이터 분포의 범위 내에 얼마나 포함되는지를 나타내는 보장 파라미터입니다. 이러한 보장은 더 강력한 가정(예: 벨만 완전성) 하에서 여러 인기 있는 알고리즘에 대해 잘 이해되어 있지만, 목표 값 함수만 특징 공간에서 선형적으로 표현 가능한 최소한의 설정에서는 이러한 이해가 부족하고 단편적입니다. 최근에는 이 설정에서의 통계적 수렴률에 대한 엄밀한 분석에 대한 관심이 높아지고 있지만, 올바른 보장의 개념은 여전히 명확하지 않으며, 이전 분석에서 제안된 후보 정의는 바람직하지 않은 특성을 가지며, 기존 문헌의 표준적인 정의와는 크게 동떨어져 있습니다. 본 논문에서는 이 설정에 대한 표준 알고리즘인 LSTDQ의 새로운 유한 표본 분석을 제시합니다. 도구 변수(instrumental variable) 관점에서 영감을 받아, 새로운 보장 파라미터인 특징-동역학 보장(feature-dynamics coverage)에 의존하는 오차 경계를 개발합니다. 특징 진화에 대한 유도된 동적 시스템에서 선형 보장으로 해석될 수 있는 이 파라미터는 추가적인 가정(예: 벨만 완전성) 하에서, 해당 설정에 특화된 보장 파라미터를 성공적으로 복원하여, 선형 OPE에서의 보장에 대한 통합적인 이해를 제공합니다.

Original Abstract

Off-policy evaluation (OPE) is a fundamental task in reinforcement learning (RL). In the classic setting of linear OPE, finite-sample guarantees often take the form $$ \textrm{Evaluation error} \le \textrm{poly}(C^π, d, 1/n,\log(1/δ)), $$ where $d$ is the dimension of the features and $C^π$ is a coverage parameter that characterizes the degree to which the visited features lie in the span of the data distribution. While such guarantees are well-understood for several popular algorithms under stronger assumptions (e.g. Bellman completeness), the understanding is lacking and fragmented in the minimal setting where only the target value function is linearly realizable in the features. Despite recent interest in tight characterizations of the statistical rate in this setting, the right notion of coverage remains unclear, and candidate definitions from prior analyses have undesirable properties and are starkly disconnected from more standard definitions in the literature. We provide a novel finite-sample analysis of a canonical algorithm for this setting, LSTDQ. Inspired by an instrumental-variable view, we develop error bounds that depend on a novel coverage parameter, the feature-dynamics coverage, which can be interpreted as linear coverage in an induced dynamical system for feature evolution. With further assumptions -- such as Bellman-completeness -- our definition successfully recovers the coverage parameters specialized to those settings, finally yielding a unified understanding for coverage in linear OPE.

2 Citations
1 Influential
5 Altmetric
29.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!