2607.27132v1 Jul 29, 2026 cs.LG

홀로노미-커버 결정 과정에서의 안정적인 몫을 이용한 최소 마르코프화

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

Zuyuan Zhang
Zuyuan Zhang
Citations: 130
h-index: 7
Tian Lan
Tian Lan
Citations: 120
h-index: 7
Mahdi Imani
Mahdi Imani
Citations: 107
h-index: 6
Yongshang Chen
Yongshang Chen
Citations: 5
h-index: 1

부분 관찰 환경에서 작동하는 에이전트는 마르코프 속성을 복원할 수 있는 재귀적으로 업데이트 가능한 통계 정보를 유지해야 하지만, 이러한 통계 정보 중 가장 작은 것은 일반적으로 알려져 있지 않습니다. 본 논문에서는 홀로노미-커버 결정 과정이라는 구조화된 POMDP 클래스에 대한 최소한의 마르코프 충분 통계량을 제시합니다. 이 클래스는 관측 가능한 동역학이 마르코프 속성을 가지며, 실제 관측 가능한 모든 전이는 숨겨진 모드에 대해 고정된 순열을 적용합니다. 특히, 본 논문에서는 일-단계 보상과 몫 상태를 유지하는 가장 세분화된 관측 기반 추상화인 '안정적인 몫(stable quotient)'을 구성하고, 현재 관측 값과 안정 클래스 쌍이 정확한 유한 마르코프 상태를 형성한다는 것을 증명합니다. 현재 클래스가 올바르게 초기화되면, 정확한 클래스 추적에는 정확히 최소 메모리 심볼이 필요하며, 이는 최대 보상 관측 시 도달 가능성 및 쌍별 결정 분리를 고려할 때 임의의 유한-메모리 제어기가 더 적은 자원을 사용할 수 없다는 의미입니다. 재설정 가능한 진단 정보를 활용하면, 가장 가까운 원형 클래스 추론은 지수적으로 감소하는 오차를 가지며, '교정 후 재시작(calibrate-then-restart)' 기법을 통해 유한-MDP의 보장을 복구된 상태로 전달할 수 있습니다. 이러한 결과는 '홀로노미 메모리 강화 학습(Holonomy Memory Reinforcement Learning)'을 가능하게 합니다. 이 방법은 현재 안정 클래스를 사용하여 메모리를 표현하고, 정렬된 엣지 전송을 통해 업데이트하며, 진단 정보가 있는 경우 로컬 클래스 좌표를 식별하고, 동기화 후에 표준 유한-MDP 강화 학습 프레임워크를 적용합니다. 실험 결과에서 원본 상태로부터 몫 상태로의 정확한 압축이 이루어지는 것을 확인했으며, 세 개의 결정 시간 메모리 상태를 사용하여 완벽한 쌍렬 순서 정확도를 달성했습니다. 이는 몫 오라클과 일치하며, 비-오라클 기반 모델보다 우수한 성능을 보입니다.

Original Abstract

An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized visible transition applies a fixed permutation to a hidden mode. In particular, we construct the stable quotient, the coarsest observation-wise abstraction preserving one-step rewards and quotient successors, and prove that the pair of the current observation and stable class forms an exact finite Markov state. When the current class is correctly initialized, exact class tracking requires exactly the minimal memory symbols, in the sense that under reachability and pairwise decision separation at a maximizing observation, no arbitrary finite-memory controller can use fewer. Under resettable diagnostics, nearest-prototype class inference has exponentially decaying error, and a calibrate-then-restart reduction transfers finite-MDP guarantees to the recovered state. The results enable \emph{Holonomy Memory Reinforcement Learning}. It represents memory by the current stable class, updates it through ordered edge transports, identifies local class coordinates when diagnostics are available, and applies a standard finite-MDP RL backbone after synchronization. Experiments recover an exact compression from raw states to quotient states and achieve perfect paired-order accuracy with three decision-time memory states, matching the quotient oracle and outperforming the non-oracle baselines.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!