2608.02197v1 Aug 03, 2026 cs.RO

중요한 곳에 집중: 시각-언어-행동 모델을 위한 적응형 시각 정제

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

Jin Cui
Jin Cui
Citations: 24
h-index: 2
Boran Zhao
Boran Zhao
Citations: 3
h-index: 1
Pengju Ren
Pengju Ren
Citations: 8
h-index: 2
Xinyue Long
Xinyue Long
Citations: 0
h-index: 0
Yanbin Hu
Yanbin Hu
Citations: 3
h-index: 1
Linkai Li
Linkai Li
Citations: 0
h-index: 0

시각-언어-행동(VLA) 모델의 시각적 표현은 공간적으로 정확한 로봇 조작에 여전히 신뢰성이 떨어집니다. 본 연구에서는 VLA 모델의 비전 인코더가 일반적인 비전 트랜스포머에서 나타나는 주의 집중 현상과 유사한 문제점을 가지고 있음을 밝히고, 또한 이러한 문제점이 훈련 후 단계에서 습득되는 공간 인식 능력과 밀접하게 관련되어 있음을 보여줍니다. 인코더가 객체 위치, 깊이 순서, 지역 기하학 정보와 같은 작업 관련 정보를 학습하는 과정에서, 제한된 글로벌 토큰 용량으로 인해 일부 정보가 낮은 정보량을 가진 패치 토큰에 유출됩니다. 본 연구에서는 AtVLA라는 프레임워크를 제안합니다. 이 프레임워크는 학습 가능한 레지스터 토큰을 비전 인코더에 삽입합니다. AtVLA는 원래의 행동 목표와 함께 로봇 데이터만을 사용하여 훈련되며, 이러한 레지스터는 몸체 인식 공간 정보를 전달하는 역할을 수행하고, 나머지 패치 토큰은 정확한 목표 위치 파악 및 미세한 접촉에 중요한 깨끗하고 공간적으로 충실한 주의 집중 분포를 회복합니다. 깨끗한 주의 집중은 신뢰할 수 있는 위치 파악을 가능하게 하지만, 저해상도 관찰에서 손실된 기하학적 세부 정보를 복구할 수는 없습니다. 따라서 AtVLA는 주의 집중 수정과 불확실성 기반의 지역 정제를 결합합니다. 행동 전문가 모델은 여러 개의 행동 단위를 샘플링하고, 이들의 불일치를 통해 불확실성을 추정합니다. 불확실한 예측에 대해서만, 행동 조건에 따른 주의 집중을 통해 작업 관련 영역을 식별하고, 잘라낸 후 고해상도로 재인코딩하여 캐시된 접두사에 추가함으로써 정교한 행동 생성이 가능합니다. LIBERO, SimplerEnv 및 실제 환경 벤치마크에서 AtVLA는 평균 LIBERO 성공률을 94.2%에서 98.4%로, 실제 환경에서의 성공률을 46.5%에서 69.0%로 향상시켰습니다. 잘라내기 작업은 전체 재계획 단계의 약 30%에서 수행되며, 대표적인 배포 설정 하에서 기본 모델보다 총 계산량이 1.4~1.6배 증가합니다.

Original Abstract

Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previously documented in generic Vision Transformers, and further show that, in embodied policies, these artifacts are closely associated with spatial perception capabilities acquired during post-training. As the encoder learns task-relevant information such as object location, depth ordering, and local geometry, limited global-token capacity causes part of this information to spill into low-information patch tokens. We introduce AtVLA, a framework that inserts learnable register tokens into the visual encoder. Trained end-to-end using only embodied data and the original action objective, these registers emerge as dedicated carriers of embodied spatial information, while the remaining patch tokens recover clean and spatially faithful attention distributions crucial for precise target localization and fine-grained contact. Clean attention restores reliable localization, but cannot recover geometric details lost in low-resolution observations. AtVLA therefore couples attention rectification with uncertainty-gated local refinement. The action expert samples multiple action chunks and estimates uncertainty from their disagreement; only for uncertain predictions, action-conditioned attention rollout identifies the task-relevant region, which is cropped, re-encoded at high resolution, and appended to the cached prefix for refined action generation. Across LIBERO, SimplerEnv, and a challenging single-view real-world benchmark, AtVLA improves the average LIBERO success rate from 94.2% to 98.4% and real-world success from 46.5% to 69.0%. The cropping is triggered on approximately 30% of replanning steps, resulting in only 1.4-1.6x the total computation of the base model under the representative deployment setting.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!