2607.08182v1 Jul 09, 2026 cs.CV

LEEVLA: 비전-언어-행동 시스템에서 잠재 환경 진화를 통해 중요한 정보를 파악하는 방법

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Xudong Wang
Xudong Wang
Citations: 46
h-index: 3
Baicheng Liu
Baicheng Liu
Citations: 46
h-index: 5
Qi Lyu
Qi Lyu
Citations: 8
h-index: 2
Lianqing Liu
Lianqing Liu
Citations: 12
h-index: 3
Zhi Han
Zhi Han
Citations: 13
h-index: 3
Jiahua Dong
Jiahua Dong
Citations: 39
h-index: 3

비전-언어-행동(VLA) 모델은 다양한 입력 데이터를 로봇의 행동으로 연결하는 것을 목표로 합니다. 그러나 기존 방식들은 시각적 토큰을 모두 동일하게 취급하고, 사람이 선택한 요소를 기반으로 추론하기 때문에 복잡한 동적인 환경에 대한 처리가 어렵고, 중요한 증거를 강조하거나 잠재적인 요인을 간과하는 문제가 있습니다. 이러한 문제를 해결하기 위해, 우리는 잠재 환경의 변화 속에서 중요한 정보를 파악하도록 설계된 VLA 아키텍처인 LEEVLA를 제안합니다. LEEVLA는 모델이 유용한 영역에 집중하도록 명시적으로 안내하고, 동시에 잠재적 세계 표현의 구조적인 변화를 유지합니다. 중요한 영역과 지시에 관련된 영역을 식별하기 위해, 우리는 동적 위치 우선순위 지정(DPP)과 의미론적 드리프트 가이드(SDG)를 결합한 드리프트 기반 동적 우선순위 지정(DGDP)을 도입하여 학습 과정에서 VLA 에이전트가 어디에 집중해야 하는지 안내합니다. 또한, 우리는 프로토타입-주변(P2P) 예측을 통해 이러한 우선순위가 부여된 특징들이 잠재 공간에서 어떻게 변화해야 하는지를 모델링하는 구조화된 특징 흐름 생성(SFFG)과, 이웃 간의 위상적 일관성을 유지하기 위한 상호 이웃 대비 손실(MC)을 도입합니다. DGDP와 SFFG는 함께 VLA 학습에 필요한 '어디'와 '어떻게'라는 측면을 고려하는 프레임워크를 구성합니다. VLA 벤치마크에서의 광범위한 실험 결과, LEEVLA는 기존 방법보다 일관되게 뛰어난 성능을 보였으며, 이는 명시적인 작업 관련 증거 가이드와 구조화된 잠재적 추론이 확장 가능한 VLA 시스템 개발에 모두 매우 중요하다는 것을 확인해 줍니다. 저희의 코드는 https://github.com/LyuQi127/LEEVLA 에서 확인할 수 있습니다.

Original Abstract

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors. To address this issue, we propose LEEVLA, a VLA architecture for seeing what matters in Latent Environment Evolution that explicitly guides the model toward informative regions while preserving the structured evolution of latent world representations. To identify salient and instruction-relevant regions, we introduce drift-guided dynamic prioritization (DGDP), which combines dynamic position prioritization (DPP) with semantic drift guidance (SDG) to guide the VLA agent where to attend during training. On top of this, we introduce structured feature flow generation (SFFG), which models how these prioritized features should evolve in latent space via prototype-to-periphery (P2P) prediction, and a mutual-neighborhood contrastive (MC) loss to maintain topological consistency among neighborhoods. Together, DGDP and SFFG form a task-aware "where-how" training framework. Extensive experiments on VLA benchmarks show that LEEVLA consistently outperforms prior methods, confirming that explicit task-evidence guidance and structured latent reasoning are both crucial for scalable VLA. Our code is available at https://github.com/LyuQi127/LEEVLA.

0 Citations
0 Influential
25.9657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!