2605.29360v1 May 28, 2026 cs.AI

MiraBench: 로봇 세계 모델의 동작 조건부 신뢰성 평가

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

Jiayi Zhou
Jiayi Zhou
Citations: 1,084
h-index: 12
Juntao Dai
Juntao Dai
Citations: 20
h-index: 3
Jiawei Chen
Jiawei Chen
Citations: 51
h-index: 3
Tianzhuo Yang
Tianzhuo Yang
Citations: 6
h-index: 1
Jiaming Ji
Jiaming Ji
Citations: 1,019
h-index: 18
Yaodong Yang
Yaodong Yang
Citations: 26
h-index: 2
Zirui Mi
Zirui Mi
Citations: 0
h-index: 0
Zhaoyi Zhang
Zhaoyi Zhang
Citations: 9
h-index: 1
Boyuan Chen
Boyuan Chen
Citations: 22
h-index: 1
Zihan Shen
Zihan Shen
Citations: 108
h-index: 5

동작 조건부 세계 모델은 로봇 학습을 위한 확장 가능한 시뮬레이터로 점점 더 많이 사용되고 있지만, 현재의 평가는 이러한 모델들의 예측이 자신이 기반으로 하는 동작 하에서 얼마나 신뢰할 수 있는지에 대한 제한적인 증거만을 제공합니다. 기존 벤치마크는 주로 시각적 충실도에 중점을 두어, 예측된 미래가 물리적으로 타당한지, 명령된 동작을 정확하게 반영하는지, 그리고 성공할 가능성이 없는 동작에 대해 얼마나 적절하게 실패를 예측하는지에 대한 명확성을 제공하지 못합니다. 본 논문에서는 로봇 세계 모델의 핵심 평가 목표인 '동작 조건부 신뢰성'을 정의하는 계층적 벤치마크인 MiraBench를 소개합니다. MiraBench는 이 목표를 세 가지 단계로 나눕니다: 물리적 일관성을 평가하는 '물리 준수 (Physics Adherence)', 작업 관련 동작 입력을 얼마나 잘 반영하는지 측정하는 '동작 충실도 (Action-Following Fidelity)', 그리고 실패를 유발하는 동작 하에서 성공적인 결과를 예측하는 경향을 파악하는 '낙관주의 편향 감지 (Optimism Bias Detection)'. 이러한 평가를 지원하기 위해, 우리는 작업, 실패 유형 및 주요 세계 모델에 대한 16,000건 이상의 판단이 포함된 인간이 직접 작성한 데이터셋을 구축했습니다. 벡터 기반 로봇 세계 모델, 텍스트 기반 생성 세계 모델, 공개 가중치 시스템, 비공개 시스템 및 다양한 모델 크기를 포괄하는 12개의 대표적인 모델 구성을 평가했습니다. 광범위한 모델 환경에서 MiraBench는 세 가지 주요 결과를 보여줍니다: 시각적 충실도는 동작 충실도의 나쁜 지표이며, 모델 크기가 증가한다고 해서 항상 동작 준수가 향상되는 것은 아니며, 현재 시스템 전반에 걸쳐 낙관주의 편향이 널리 퍼져 있습니다. MiraBench는 평가 기준을 외형에서 동작 조건부 신뢰성으로 전환함으로써, 로봇 세계 모델을 충실한 시뮬레이터로 평가하고 개선하기 위한 기반을 제공합니다.

Original Abstract

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias Detection}, which probes the tendency to predict successful outcomes under failure-inducing actions. To support this evaluation, we curate a human-annotated corpus with over 16,000 judgments across tasks, failure categories, and leading world models. We evaluate 12 representative model configurations spanning vector-conditioned robotic world models, text-conditioned generative world models, open-weight systems, closed-source systems, and multiple model scales. Across this broad model landscape, MiraBench reveals three central findings: visual fidelity is a poor proxy for action fidelity; increasing model scale does not reliably improve action following; and optimism bias is pervasive across current systems. By shifting evaluation from appearance to action-conditioned reliability, MiraBench provides a diagnostic foundation for assessing and improving robotic world models as faithful simulators.

1 Citations
0 Influential
9 Altmetric
46.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!