MotionHalluc: 미세 수준의 동작 추론에서 발생하는 운동 환각 진단
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
교차 동영상 비교에서 동작 명령 생성은 쿼리 동영상과 참조 동영상 간의 차이점을 설명하는 교정 피드백을 생성하는 것을 목표로 합니다. 그러나 기존 모델은 종종 실제 동영상 쌍 간의 운동학적 차이를 반영하지 못하고 운동 환각 현상을 보이는 명령어를 생성하는 경우가 많습니다. 이러한 환각 현상을 체계적으로 조사하기 위해, 우리는 paired-video 비교에서 운동 환각을 평가하기 위한 전용 벤치마크인 MotionHalluc을 소개합니다. MotionHalluc은 553개의 동영상 쌍에 대한 1540개의 미세 수준 질문으로 구성되어 있으며, 다음과 같은 세 가지 핵심 측면에서 환각 현상을 평가합니다: (1) 방향성 환각, (2) 귀속성 환각, 그리고 (3) 시간적 환각. 최첨단 대규모 멀티모달 모델에 대한 광범위한 평가는 이러한 환각 현상에 매우 취약하다는 것을 보여줍니다. 또한, 우리는 훈련 과정 없이 측정값을 추출하고 검증하는 baseline인 Perceive-Parse-Verify (PPV)를 제공합니다. PPV는 후보 명령어를 실행 가능한 측정 쿼리로 변환하고 추론 시점에 운동학적 측정을 제공합니다. 우리의 결과는 이 간단한 측정값 주입 방식이 모델 전반에 걸쳐 평균 10.6%의 성능 향상을 가져온다는 것을 보여주며, 이는 명시적인 정량적 측정을 사용한 동작 추론이 교차 동영상 비교에서 환각 현상을 줄이는 데 중요한 요소임을 시사합니다. 본 논문 채택 시 코드와 데이터셋을 공개할 예정입니다.
Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a reference motion. However, existing models often generate instructions that exhibit motion hallucinations, failing to reflect actual kinematic differences between paired videos. To systematically investigate these hallucinations, we introduce MotionHalluc, a dedicated benchmark for evaluating motion hallucinations in paired-video comparison. MotionHalluc comprises 1540 fine-grained questions over 553 video pairs, evaluating hallucinations along three core dimensions: (1)directional hallucination, (2)attributional hallucination, and (3)temporal hallucination. Extensive evaluations of state-of-the-art large multimodal models demonstrate high susceptibility to these hallucinations. Furthermore, we provide Perceive-Parse-Verify (PPV) as a training-free measurements extraction and verification baseline that converts candidate instructions into executable measurement queries and supplies kinematic measurements at inference time. Our results show that this simple measurements injection yields an average 10.6% performance gain across models, suggesting that motion reasoning with explicit quantitative measurements is a key factor in reducing hallucinations in cross-video comparison. Our code and dataset will be made publicly available upon acceptance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.