MPIE-Bench: 해부학적으로 타당한 다인체 상호작용 편집 성능 평가
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
최근의 텍스트-이미지 및 개인 맞춤형 편집 모델은 고품질 단일 인물 이미지를 쉽게 생성할 수 있게 되었습니다. 그러나 여러 명의 특정 인물을 포용, 운반 또는 격투와 같은 공통적인 상호작용 상황에 배치하는 것은 여전히 심각한 문제를 드러냅니다. 예를 들어, 융합된 팔다리, 존재하지 않는 신체 부위, 서로 겹쳐진 몸 등이 발생합니다. 기존의 평가지표는 이러한 해부학적 및 기하학적 문제들을 충분히 고려하지 못하며, VLM(Vision-Language Model) 기반 평가 항목은 종종 상호작용 측면에서만 높은 점수를 나타내지만, 인간에게는 명백한 오류가 존재합니다. 우리는 405개의 장면, 14가지 상호작용 유형 및 네 가지 밀도(C0-C3)를 포괄하는 2,500개 샘플로 구성된 비디오 기반 편집 데이터셋인 MPIE-Bench를 소개합니다. 또한, 공개된 다인체 메시 재구성 데이터를 활용하여 접촉 시간의 기하학적 정확성을 평가하는 새로운 지표인 MPIE-Eval을 제안합니다. '해부학(Anatomy)'은 모든 인간과 유사한 덩어리가 완전하게 재구현된 신체로 설명될 수 있는지 여부를 평가하며, '상호작용(Interaction)'은 해당 신체 간의 침투 및 표면 거리가 지시된 상호작용에 부합하는지 여부를 평가합니다. 10개의 편집 모델을 테스트한 결과, 메시 기반 해부학 정확도는 최대 0.65이고, 메시 기반 상호작용 정확도는 최대 0.72로 나타났습니다. 즉, 어떤 단일 모델도 두 가지 측면 모두에서 뛰어난 성능을 보이지 않습니다. 반면, VLM 체크리스트는 동일한 이미지에 대해 0.95 이상의 높은 점수를 부여합니다. 다섯 명의 평가자가 참여한 연구 결과, 제안된 두 가지 지표가 VLM 기반 평가보다 인간의 판단과 더 밀접하게 일치하며, 모든 가중치와 임계값을 제거해도 이러한 순위는 유지됨을 확인했습니다.
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.