2607.27508v1 Jul 29, 2026 cs.RO

단일 라운드에서의 수정 가능한 지원: 실용적-교육적 최적 반응

Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

J. F. Fisac
J. F. Fisac
Citations: 151
h-index: 8
Elle Lazarski
Elle Lazarski
Citations: 11
h-index: 1

지원 게임은 비대칭 정보 하에서 인간과 로봇의 협력을 모델링합니다. 여기서 인간은 목표를 알고 있지만, 로봇은 효과적인 지원을 위해 관찰 및 상호 작용을 통해 이를 추론해야 합니다. 일반적으로 최적의 지원 게임 전략을 실시간으로 계산하는 것은 매우 어렵습니다. 왜냐하면 정확한 해를 구하려면 POMDP(부분 관측 마르코프 결정 프로세스)에서의 계획이 필요하기 때문입니다. 우리는 특정 유형의 지원 게임에서, 실용적-교육적 추론이 단일 시간 단계 내에서 목표 불확실성을 해결하여 전체 시계열 게임을 처리 가능한 최적 반응 절차를 통해 정확하게 풀 수 있음을 밝힙니다. 이러한 유형 내에서, 널리 사용되는 역방향 최적 제어 방법은 정렬을 방해하는 추론의 한계를 가지고 있지만, 실용적-교육적 추론은 작업 실행만으로는 동일해 보이는 행동을 통해 즉시 목표를 명확하게 하여 이 장벽을 극복합니다. 마지막으로, 우리는 간단한 협업 블록 조립 예제를 사용하여 이론적인 결과를 검증하고 제안된 방법을 평가했습니다.

Original Abstract

Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effectively. In general, computing optimal assistance game strategies online is intractable, since exact solutions require planning in a POMDP. We identify a class of assistance games in which pragmatic-pedagogic reasoning resolves goal uncertainty in a single time step, rendering the full-horizon game exactly solvable by a tractable best-response procedure. Within this class, we show that mainstream inverse optimal control exhibits an inference ceiling that hinders alignment, while pragmatic-pedagogic reasoning overcomes this barrier by immediately disambiguating goals through actions that look equivalent under task execution alone. Finally, we validate our theoretical results and proposed method on a simple collaborative block-building example.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!