2606.10366v1 Jun 09, 2026 cs.RO

VLA 평가를 위한 시뮬레이션과 실제 환경 간의 상관관계 개선을 위한 실용적인 방법

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation

Yingdong Hu
Yingdong Hu
Citations: 974
h-index: 13
Shuo Wang
Shuo Wang
Citations: 104
h-index: 3
Hanyuan Xu
Hanyuan Xu
Citations: 34
h-index: 2
Fanqi Lin
Fanqi Lin
Citations: 763
h-index: 7
Yang Gao
Yang Gao
Citations: 17
h-index: 2

시뮬레이션은 비전-언어-행동(VLA) 정책을 평가하고 개선하는 데 필수적인 도구로 자리 잡았습니다. 이는 비용이 많이 드는 실제 로봇 평가에 대한 확장 가능하고, 재현 가능하며, 제어 가능한 대안을 제공합니다. 최근 시뮬레이션 벤치마크는 현실성과 다양성 측면에서 상당한 발전을 이루었지만, 이러한 플랫폼은 여전히 실제 환경에서의 정책 평가를 위한 신뢰할 수 있는 지표로 널리 채택되지 못하고 있습니다. 본 연구에서는 시뮬레이션과 실제 환경 간의 상관관계라는 관점에서 이 문제를 조사합니다. 우리는 여러 시뮬레이션 플랫폼, VLA 정책, 작업 및 교란 요인에 대한 체계적인 연구를 수행하여 시뮬레이션 평가가 정책 순위 일관성, 성능 상관 관계 및 교란 요인별 실패 패턴 측면에서 실제 환경에서의 결론을 유지하는지 측정합니다. 이러한 분석을 통해 기존 시뮬레이터의 한계를 파악하고 어떤 종류의 시뮬레이션 신호가 실제 배포와 더 잘 일치하는지 밝혀낼 수 있습니다. 또한, 사용자가 정책 개선을 위해 시뮬레이션을 어떻게 활용해야 하는지에 대해 연구합니다. 여기에는 시뮬레이터를 기반으로 한 미세 조정이 언제 유익한지, 그리고 훈련 후 데이터의 양이 시뮬레이션과 실제 환경 간의 일치성에 어떤 영향을 미치는지에 대한 내용이 포함됩니다. 전반적으로 본 연구는 VLA 정책을 위한 시뮬레이션의 유용성을 측정, 해석 및 개선하기 위한 통합 프레임워크를 제공하며, 이는 시뮬레이터 설계자와 정책 개발 파이프라인의 일부로 시뮬레이션을 사용하는 실무자 모두에게 지침을 제공합니다.

Original Abstract

Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and controllable alternatives to costly real-world robot evaluation. Recent simulation benchmarks have made substantial progress on realism and diversity, yet these platforms have not been widely adopted as reliable proxies for real-world policy evaluation. In this work, we investigate this issue through the lens of sim-and-real correlation. We conduct a systematic study across multiple simulation platforms, VLA policies, tasks, and perturbation factors, measuring whether simulated evaluation preserves real-world conclusions in terms of policy ranking consistency, performance correlation, and perturbation-wise failure patterns. This analysis allows us to characterize the limitations of existing simulators and identify what kinds of simulation signals are more aligned with real-world deployment. We further examine how users should exploit simulation for policy improvement, including when simulator-based finetuning is beneficial and how the amount of post-training data affects sim-and-real alignment. Overall, our work provides a unified framework for measuring, interpreting, and improving the usefulness of simulation for VLA policies, offering guidance both for simulator designers and for practitioners who use simulation as part of the policy development pipeline.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!