2607.29622v1 Jul 31, 2026 cs.RO

RayViT: 시점 변화에 강건한 모방 학습을 위한 광선 기반 시각적 표현

RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

N. Freymuth
N. Freymuth
Citations: 244
h-index: 8
Ge Li
Ge Li
Citations: 279
h-index: 10
Qian Wang
Qian Wang
Citations: 79
h-index: 5
Weiran Liao
Weiran Liao
Citations: 35
h-index: 3
Longrui Chen
Longrui Chen
Citations: 4
h-index: 1
Peiran Sun
Peiran Sun
Citations: 0
h-index: 0
Aleksandar Taranovic
Aleksandar Taranovic
Citations: 61
h-index: 5
C. F. M. Nagy
C. F. M. Nagy
Citations: 0
h-index: 0
Y. Tan
Y. Tan
Citations: 0
h-index: 0
Gerhard Neumann
Gerhard Neumann
Citations: 308
h-index: 10
Tao Chen
Tao Chen
Citations: 0
h-index: 0

시각적 모방 학습은 로봇이 이미지로부터 직접 시지각-운동 기술을 습득하도록 하지만, RGB 관찰 데이터는 명시적인 기하학적 정보를 포함하지 않아 학습된 정책이 카메라 변화에 취약해지는 문제가 있습니다. 이를 해결하기 위해, 본 연구에서는 사전 학습된 ViT 모델의 기반 구조에 카메라 기하학 정보를 주입하는 경량화된 아키텍처인 extbf{Ray-conditioned Vision Transformer Encoder (RayViT)}를 제안합니다. RayViT는 카메라 기하학 정보를 플뤼커 광선 맵으로 표현하고, 이를 광선 특징으로 분할하여 게이티드 크로스 어텐션을 사용하여 광선 기반 클래스 토큰을 생성합니다. 이러한 광선 특징은 밀집된 위치 임베딩으로 추가되고, 광선 클래스 토큰은 원래 ViT 클래스 토큰을 대체하여 기하학 정보를 고려한 요약 표현을 제공합니다. 또한, 이 접근 방식을 보조 코사인 유사성 손실과 결합하여 기하학 정보에 대한 토큰의 성능과 강건성을 지속적으로 향상시킵니다. 시뮬레이션 및 실제 로봇 환경에서의 실험 결과는 RayViT가 멀티 태스크 RoboCasa 벤치마크에서 카메라 변화에 따른 강건성이 약 13% 포인트 향상되고, 실제 환경에서의 멀티 태스크 성공률이 평균 1.78개의 단계를 추가로 완료한다는 것을 보여줍니다.

Original Abstract

Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making learned policies brittle to camera perturbations. To address this, we propose \textbf{Ray-conditioned Vision Transformer Encoder (RayViT)}, a lightweight architecture that injects camera geometry into pretrained ViT backbones. RayViT represents camera geometry as a Plücker ray map, patchifies it into ray features, and uses gated cross-attention to produce a ray-conditioned class token. These ray features are added as dense positional embeddings, while the ray class token replaces the original ViT class token to provide a geometry-aware summary representation. We combine this approach with an auxiliary cosine similarity loss to consistently improve the performance and robustness for geometry-aware tokens. Experiments on sim- and real-robot tasks demonstrate that RayViT improves robustness by approximately 13 percentage points under camera perturbations in multi-task RoboCasa benchmark and by 1.78 average completed stages in real-world multi-task success rate compared to baselines.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!