2607.08647v1 Jul 09, 2026 cs.LG

강건한 보상 학습을 위한 다중 모드, 다중 환경 머신 티칭

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

Daniel S. Brown
Daniel S. Brown
Citations: 2
h-index: 1
Ali Larian
Ali Larian
Citations: 5
h-index: 1
Qian Lin
Qian Lin
Citations: 56
h-index: 4
Changzhong Wu
Changzhong Wu
Citations: 0
h-index: 0

자율 에이전트가 다양한 운영 환경에 점점 더 많이 배치됨에 따라, 인간의 의도와 일치하는 에이전트 행동은 특정 환경에 과적합되는 것이 아니라 그러한 변화에도 강건하게 유지될 수 있는 보상 함수를 요구합니다. 역강화학습(IRL)은 인간 피드백으로부터 이러한 목표를 추론하는 체계적인 방법을 제공합니다. 그러나, 최적의 IRL 티칭 접근 방식에 대한 기존 연구는 단일 환경 및 데모 데이터만을 사용하는 설정에 초점을 맞추고 있으며, 다양한 피드백 모달리티와 환경 동역학이 여러 환경에 걸쳐 일반화되는 보상 함수를 어떻게 공동으로 제약하는지에 대한 탐구가 부족합니다. 왜냐하면 하나의 MDP(Markov Decision Process)에서의 데모는 보상 정보를 해당 환경의 특정 구조와 얽히게 하기 때문에, 결과적으로 생성된 보상은 에이전트가 새로운 환경에 배치될 때 자주 일반화되지 않습니다. 본 연구에서는 먼저 다양한 피드백 모달리티가 보상을 어떻게 제약하는지 분석하고, 무한 데이터 영역에서 비교(comparison)는 다른 모달리티보다 훨씬 강력한 전역적 제약을 가한다는 것을 보여줍니다. 이러한 이론적 분석 외에도, 여러 MDP 환경에서 작동하는 계층적 머신 티칭 알고리즘을 소개합니다. 이 알고리즘은 먼저 상호 보완적인 보상 제약을 드러내는 유용한 환경을 탐욕적으로 선택하고, 그런 다음 해당 환경 내에서 저렴한 피드백을 전략적으로 요청합니다. 실험 결과, 본 방법은 동일한 피드백 예산 하에 균일한 티칭 기준선보다 훨씬 낮은 후회(regret)를 달성하고 보류된 환경으로 더 강력하게 일반화되는 것을 보여주며, 이는 다중 환경, 다중 모드 티칭이 동역학적으로 강건한 보상 함수를 학습하는 데 얼마나 중요한지를 입증합니다.

Original Abstract

As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how heterogeneous feedback modalities and environment dynamics jointly constrain reward functions that generalize across multiple environments. Because demonstrations in one MDP entangle reward information with that environments specific structure, the resulting rewards frequently fail to generalize when the agent is deployed in a new setting. We first analyze how different feedback modalities constrain rewards, showing that, in the unlimited-data regime, comparisons impose strictly stronger global constraints than other modalities. Beyond this theoretical analysis, we introduce a hierarchical machine teaching algorithm for reward learning that operates across multiple MDPs. The algorithm first greedily selects informative environments that expose complementary reward constraints, then strategically queries low-cost feedback within those environments. Empirically, our method achieves substantially lower regret and stronger generalization to held-out environments than uniform teaching baselines under identical feedback budgets, demonstrating the importance of multi-environment, multi-modal teaching for learning dynamics-robust reward functions.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!