CheckVLA: 액션-조건부 세계 모델을 활용한 장기 모바일 조작의 실행 시간 검증
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
비전-언어-액션(VLA) 정책은 일반적으로 개방 루프 액션 단위로 장기 모바일 조작을 수행하며, 새로운 고수준 시각적 입력을 받지 않고 여러 개의 액션을 실행합니다. 한 번 실행된 액션 단위를 '커밋'이라고 할 때, 이 커밋은 관측값이 어떻게 변화해야 하는지를 나타내지만, 의도치 않은 편차는 이러한 기대를 위반할 수 있으며, 나머지 액션들은 오류를 계속 전파합니다. 정책의 신뢰도는 실행 시점에 발생한 편차에 대응할 수 없으며, 관측값만을 사용하는 이상 감지 방법은 예상되는 효과와 설명되지 않는 변화를 구별하기 위한 액션-조건부 기준점을 제공하지 못합니다. 본 논문에서는 별도로 학습되고 고정된 액션-조건부 세계 모델을 사용하여 실행을 검증하는 CheckVLA를 제안합니다. 컨포멀 방식으로 조정된 위험 임계값은 불필요한 초기 개입의 에피소드 수준 확률을 제한하고, 언제 개입해야 하는지를 결정하며, 이 임계값을 초과하면 수정된 후반부가 이전 액션 단위를 얼마나 유지하는지에 대한 제어를 제공합니다. 지연 시간을 고려한 하드 프리픽싱은 대체 가능한 액션으로만 대체하도록 제한하며, 이벤트 기반 키프레임 저장소는 수리 과정에서 이전 진행 상황에 대한 증거를 보존합니다. RoboCasa365 환경에서 동일한 학습 레시피와 호출 예산을 사용했을 때, CheckVLA는 평균 36.1%의 성공률을 달성하여 주기적인 재계획 방법을 사용하는 경우보다 8.5% 포인트 더 높은 성능을 보였습니다. 동일한 5% 에피소드 수준의 오탐율 목표에서, 액션 조건부는 시기 적절한 복구율을 77.9%로 높여 관측값만 사용하는 제어 방법의 48.6%와 액션을 무작위로 섞는 제어 방법의 37.9%보다 훨씬 우수한 성능을 보였습니다. 이러한 시뮬레이션 결과는 액션-조건부 검증이 덩어리 단위 실행 중에 피드백을 복원하는 동시에 수리 과정이 추론 지연 시간과 일관성을 유지하도록 하는 효과적인 방법임을 보여줍니다.
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.