2606.10862v1 Jun 09, 2026 cs.CV

LIBERO-Occ: 시야 가림 현상으로 인한 문제점을 평가하고 개선하는 Vision-Language-Action 모델 연구 - 관점 상상 기법 활용

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination

Jiwen Zhang
Jiwen Zhang
Citations: 249
h-index: 5
Siyuan Wang
Siyuan Wang
Citations: 1,445
h-index: 21
Zhongyu Wei
Zhongyu Wei
Citations: 347
h-index: 6
Taishan Li
Taishan Li
Citations: 100
h-index: 6
Xuanjing Huang
Xuanjing Huang
Citations: 712
h-index: 13

Vision-Language-Action (VLA) 모델은 표준 조작 벤치마크에서 뛰어난 성능을 보이지만, 대부분의 평가는 작업과 관련된 객체가 완전히 보이는 상태를 가정합니다. 이러한 가정은 실제 환경에서 종종 실패하며, 시야 가림 현상은 조작 과정을 부분적으로 관찰하기 어렵게 만듭니다. 본 논문에서는 VLA 모델에 대한 근본적인 과제로 *시야 가림 현상*을 연구하고, LIBERO의 시야 가림 현상 지향 확장 버전인 extbf{LIBERO-Occ}를 소개합니다. 실험 결과, 최첨단 VLA 모델들이 시야 가림 현상이 발생할 경우 상당한 성능 저하를 겪는 것으로 나타났습니다. 이러한 문제를 해결하기 위해, 본 연구에서는 extbf{관점 상상 (VIM)} 기법을 제안합니다. VIM은 가려진 객체에 대한 보조적인 시점을 생성하고, 관찰된 정보와 상상된 정보를 모두 활용하여 행동 예측을 수행합니다. VIM은 추가 카메라 없이도 다양한 작업 환경, 시야 가림 유형 및 심각도 수준에서 모델의 안정성을 향상시키며, 부분적으로 보이는 조작 환경에서의 인지 능력 보완에 있어 유망한 기술임을 보여줍니다. 본 연구의 벤치마크 데이터와 관련 코드는 다음 링크에서 확인할 수 있습니다: [https://github.com/litsh/Libero-Occ](https://github.com/litsh/Libero-Occ)

Original Abstract

Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible. This assumption often fails in realistic settings, where occlusion makes manipulation partially observable. In this paper, we study \textit{scene-induced occlusion} as a fundamental challenge for VLA models and introduce \textbf{LIBERO-Occ}, an occlusion-oriented extension of LIBERO. Experiments show that state-of-the-art VLAs suffer substantial performance degradation under occlusion. To address this issue, we propose \textbf{Viewpoint Imagination (VIM)}, which generates a complementary view from an occluded primary observation and conditions action prediction on both observed and imagined evidence. VIM improves robustness across task suites, occlusion types, and severity levels without requiring additional cameras at deployment time, suggesting that viewpoint imagination is an promising mechanism for perception completion in partially observable manipulation. Our benchmark and corresponding code are available at: \href{https://github.com/litsh/Libero-Occ}{https://github.com/litsh/Libero-Occ}.

0 Citations
0 Influential
30.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!