2604.05673v1 Apr 07, 2026 cs.RO

적정화된 슈뢰딩거 브리지 매칭을 이용한 소규모 단계 시각적 내비게이션

Rectified Schrödinger Bridge Matching for Few-Step Visual Navigation

Wenjian Zhang
Wenjian Zhang
Citations: 1,666
h-index: 4
Wuyang Luan
Wuyang Luan
Citations: 12
h-index: 2
Junhui Li
Junhui Li
Citations: 14
h-index: 2
Rui Ma
Rui Ma
Citations: 5
h-index: 1
Weiguang Zhao
Weiguang Zhao
Citations: 120
h-index: 5
Tieru Wu
Tieru Wu
Citations: 28
h-index: 3

시각적 내비게이션은 자율 에이전트가 고차원 감각 정보를 연속적이고 장기적인 행동 경로로 변환해야 하는, 임베디드 AI의 핵심적인 과제입니다. 확산 모델 및 슈뢰딩거 브리지(SB)를 기반으로 하는 생성 정책은 다양한 행동 분포를 효과적으로 모델링하지만, 높은 분산의 확률적 변환으로 인해 수십 번의 적분 단계가 필요하며, 이는 실시간 로봇 제어에 중요한 장애물이 됩니다. 본 논문에서는 표준 슈뢰딩거 브리지(ε=1, 최대 엔트로피 변환)와 결정론적 최적 변환(ε→0, Conditional Flow Matching과 유사) 사이의 공통 속도장 구조를 활용하는 Rectified Schrödinger Bridge Matching (RSBM) 프레임워크를 제안합니다. 이 프레임워크는 단일 엔트로피 정규화 파라미터 ε에 의해 제어됩니다. 우리는 두 가지 중요한 결과를 증명합니다. (1) 조건부 속도장의 함수 형태는 전체 ε 스펙트럼에 걸쳐 불변이며(Velocity Structure Invariance), 이를 통해 단일 네트워크가 모든 정규화 강도를 처리할 수 있습니다. (2) ε 값을 선형적으로 감소시키면 조건부 속도장의 분산이 감소하여 보다 안정적인 초기 단계 ODE 적분을 가능하게 합니다. 학습된 조건부 사전 지식을 활용하여 변환 거리를 단축하는 RSBM은 다양한 분포를 포괄하고 경로의 직선성을 균형 있게 유지하는 중간 ε 값을 사용합니다. 실험적으로, 표준 브리지는 수렴하는 데 ≥ 10단계가 필요하지만, RSBM은 별도의 증류 또는 다단계 훈련 없이 단 3단계 적분만으로 94% 이상의 코사인 유사도와 92%의 성공률을 달성합니다. 이는 고정밀 생성 정책과 임베디드 AI의 낮은 지연 요구 사항 간의 격차를 크게 줄여줍니다.

Original Abstract

Visual navigation is a core challenge in Embodied AI, requiring autonomous agents to translate high-dimensional sensory observations into continuous, long-horizon action trajectories. While generative policies based on diffusion models and Schrödinger Bridges (SB) effectively capture multimodal action distributions, they require dozens of integration steps due to high-variance stochastic transport, posing a critical barrier for real-time robotic control. We propose Rectified Schrödinger Bridge Matching (RSBM), a framework that exploits a shared velocity-field structure between standard Schrödinger Bridges ($\varepsilon=1$, maximum-entropy transport) and deterministic Optimal Transport ($\varepsilon\to 0$, as in Conditional Flow Matching), controlled by a single entropic regularization parameter $\varepsilon$. We prove two key results: (1) the conditional velocity field's functional form is invariant across the entire $\varepsilon$-spectrum (Velocity Structure Invariance), enabling a single network to serve all regularization strengths; and (2) reducing $\varepsilon$ linearly decreases the conditional velocity variance, enabling more stable coarse-step ODE integration. Anchored to a learned conditional prior that shortens transport distance, RSBM operates at an intermediate $\varepsilon$ that balances multimodal coverage and path straightness. Empirically, while standard bridges require $\geq 10$ steps to converge, RSBM achieves over 94% cosine similarity and 92% success rate in merely 3 integration steps -- without distillation or multi-stage training -- substantially narrowing the gap between high-fidelity generative policies and the low-latency demands of Embodied AI.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!