시간적 추상화를 통한 순방향-역방향 표현에서의 스펙트럼 정렬
Spectral Alignment in Forward-Backward Representations via Temporal Abstraction
순방향-역방향(FB) 표현은 저랭크 분해를 통해 연속 공간에서 서세션 표현(SR)을 학습하는 강력한 프레임워크를 제공합니다. 그러나 연속 환경의 고랭크 전이 동역학과 FB 아키텍처의 저랭크 병목 현상 사이에 근본적인 스펙트럼 불일치가 종종 존재하며, 이는 정확한 저랭크 표현 학습을 어렵게 만듭니다. 본 연구에서는 시간적 추상화를 이러한 불일치를 완화하는 메커니즘으로 분석합니다. 전이 연산자의 스펙트럼 특성을 분석한 결과, 시간적 추상화는 고주파 스펙트럼 구성 요소를 억제하는 저역 통과 필터 역할을 한다는 것을 보여줍니다. 이러한 억제는 유도된 SR의 효과적인 랭크를 줄이는 동시에 결과 값 함수 오류에 대한 형식적인 경계를 유지합니다. 실험적으로, 이러한 정렬은 특히 부트스트래핑이 오류 발생 가능성이 높은 높은 할인 계수에서 안정적인 FB 학습에 중요한 요인임을 보여줍니다. 우리의 결과는 시간적 추상화가 기본 MDP의 스펙트럴 구조를 형성하고 연속 제어에서 효과적인 장기 표현을 가능하게 하는 원칙적인 메커니즘임을 보여줍니다.
Forward-backward (FB) representations provide a powerful framework for learning the successor representation (SR) in continuous spaces by enforcing a low-rank factorization. However, a fundamental spectral mismatch often exists between the high-rank transition dynamics of continuous environments and the low-rank bottleneck of the FB architecture, making accurate low-rank representation learning difficult. In this work, we analyze temporal abstraction as a mechanism to mitigate this mismatch. By characterizing the spectral properties of the transition operator, we show that temporal abstraction acts as a low-pass filter that suppresses high-frequency spectral components. This suppression reduces the effective rank of the induced SR while preserving a formal bound on the resulting value function error. Empirically, we show that this alignment is a key factor for stable FB learning, particularly at high discount factors where bootstrapping becomes error-prone. Our results identify temporal abstraction as a principled mechanism for shaping the spectral structure of the underlying MDP and enabling effective long-horizon representations in continuous control.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.