2608.01573v1 Aug 03, 2026 cs.RO

시각-언어-행동 모델의 위치 기반 취약점 발견 및 완화

Uncovering and Mitigating Positional Blind Spots in Vision-Language-Action Models

Qin Zhao
Qin Zhao
Citations: 0
h-index: 0
Yihao Huang
Yihao Huang
Citations: 1,097
h-index: 19
Dongdong An
Dongdong An
Citations: 58
h-index: 3
Pengjie Zhao
Pengjie Zhao
Citations: 0
h-index: 0
Wenbing Tang
Wenbing Tang
Citations: 85
h-index: 6
Ziming He
Ziming He
Citations: 46
h-index: 3
Jiayi Zhu
Jiayi Zhu
Citations: 10
h-index: 2
Jifeng Ning
Jifeng Ning
Citations: 0
h-index: 0

최근 시각-언어-행동(VLA) 모델은 로봇 조작 분야에서 뛰어난 성능을 보이며, 일반적으로 사전에 정의된 객체 배치 환경에 대한 성공률로 평가됩니다. 이러한 평가는 작업 공간 전체에 걸쳐 균일한 능력을 가정하지만, 이는 사실이 아닙니다. 연구에서는 명령과 다른 모든 장면 요소가 동일하게 유지되는 상황에서도, 관련 없는 방해물을 단순히 이동시키기만 하면 특정 영역에서 실패 확률이 크게 증가하는 현상을 '위치 기반 취약점(PBS)'이라고 명명했습니다. 본 논문에서는 PBS를 발견하고 완화하기 위한 두 단계의 블랙박스 프레임워크를 제안합니다. 첫 번째 단계인 발견 단계에서는 작업 공간을 격자로 나누고, 일방향 로그-우도 비율 검정을 사용하여 위험이 크게 증가한 PBS 영역을 식별합니다. 두 번째 단계인 완화 단계에서는 이러한 PBS 영역에서 수집된 데이터를 기반으로 LoRA(Low-Rank Adaptation)를 통해 정책을 미세 조정하여 해당 영역의 성능을 향상시키면서 나머지 작업 공간에서의 전반적인 성능 저하를 최소화합니다. 제안하는 프레임워크를 두 개의 벤치마크 환경에서 최첨단 VLA 모델 5가지에 대해 평가한 결과, 모든 모델에서 PBS가 광범위하게 나타나며 특정 영역에 집중되어 있다는 것을 확인했습니다. 또한, 제안하는 탐색 전략은 F1 점수가 0.678로, 무작위 탐색 및 적응적 샘플링 기준보다 각각 0.268과 0.178만큼 우수한 성능을 보였습니다. 발견된 영역을 기반으로 한 목표 미세 조정은 전체 실패율을 40.00%에서 85.19%까지 감소시키는 효과를 가져왔습니다.

Original Abstract

Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically measured by success rates aggregated over predefined object configurations, an evaluation that implicitly assumes spatially uniform competence across the workspace. However, this assumption does not hold: even with the instruction and every other scene factor held fixed, merely relocating a task-irrelevant distractor can sharply raise the failure probability within localized, spatially coherent regions, which we term Positional Blind Spots (PBS). In this paper, we propose a two-stage black-box framework to uncover and mitigate PBS. During the uncovering stage, we grid the workspace and apply a one-sided log-likelihood-ratio test to localize PBS cells with significantly elevated risk. During the mitigation stage, we fine-tune the policy via LoRA on demonstrations collected from these PBS regions, improving competence there while largely preserving performance across the rest of the workspace. We evaluate our framework on five state-of-the-art VLA policies across two benchmarks, and find that PBS are pervasive and spatially concentrated in all of them, with failure rates up to 0.58. Our search strategy achieves an average F1-score of 0.678, outperforming random search and adaptive sampling baselines by 0.268 and 0.178, respectively. Guided by the discovered regions, targeted fine-tuning reduces the overall failure rate by 40.00%--85.19%.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!