TriViewBench: 다중 시점 구조적 추론을 위한 복잡도 제어 및 확장성 평가
TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
다중 모달 대규모 언어 모델(MLLM)은 표준적인 시각 질의응답 벤치마크에서 뛰어난 성능을 보이지만, 제어된 구조적 복잡성 하에서의 확장성은 아직 제대로 이해되지 못하고 있습니다. 본 연구에서는 합성된 3D 장면으로 구성되고 객체 수 및 가려짐 정도가 명시적으로 설정되어 있는, 복잡도 제어를 위한 시각 추론 벤치마크인 TriViewBench를 소개합니다. 이 벤치마크는 1,923개의 장면과 14,000개 이상의 질의응답(QA) 쌍으로 구성되며, 네 가지 복잡성 수준과 세 가지 추론 범주(지역적 판단, 객체 계수, 전체적인 재구성)로 분류됩니다. 본 연구에서는 18개의 공개 및 비공개 MLLM을 통일된 프롬프트 프로토콜 하에서 평가했습니다. 모든 18개의 모델이 예외 없이 동일한 성능 계층 구조를 나타냈으며, 성능은 복잡도에 따라 단조롭게 저하되었습니다. 지역적 판단 작업의 성능 저하는 미미했지만(상대적으로 12.11% 감소), 객체 계수의 성능 저하는 상당했으며(59.14%), 전체적인 재구성은 심각하게 저하되었습니다(80.02%). 객체 계수 오류 분석 결과, 두 가지 독립적인 실패 메커니즘이 밝혀졌습니다. 단일 시점 작업에서는 가려짐으로 인한 과소 계수가 지배적이며, 다중 시점 작업에서는 서로 다른 시점에서 동일한 객체를 인식하는 데 어려움으로 인해 과대 계수가 발생합니다. Chain-of-Thought(CoT) 프롬프트는 전반적으로 거의 효과가 없었습니다($Δ = -0.16%$) 또한 CoT의 효과는 전체적인 재구성에 대해 기능적 제약을 받으며, 이는 추론 전략보다는 시점 간 공간 표현에서 병목 현상이 발생한다는 것을 시사합니다. 이러한 결과는 현재 MLLM의 근본적인 확장성 한계를 드러내며, TriViewBench를 구조적 추론 실패를 분석하기 위한 제어된 진단 프레임워크로 제시합니다.
Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexity remains poorly understood. We introduce TriViewBench, a controlled three-view visual reasoning benchmark constructed from synthetic 3D scenes with explicitly parameterized object count and occlusion. The benchmark contains 1,923 scenes and over 14K Question-Answer (QA) pairs organized into four complexity levels and three reasoning categories: Local Decision, Object Counting, and Global Recovery. We evaluate 18 open- and closed-source MLLMs under a unified prompting protocol. All 18 models exhibit an identical capability hierarchy without exception (Local Decision > Object Counting > Global Recovery), and performance degrades monotonically with complexity: Local Decision tasks decline modestly (12.11% relative drop), while Object Counting degrades substantially (59.14%) and Global Recovery collapses severely (80.02%). Error analysis on Object Counting reveals two mechanistically independent failure modes: single-view tasks are dominated by undercounting due to occlusion blindness, whereas the multi-view task reverses to overcounting due to cross-view identity confusion. Chain-of-Thought (CoT) prompting yields near-zero overall benefit ($Δ= -0.16\%$) and its effect on Global Recovery is strongly capability-gated, suggesting that the bottleneck lies in cross-view spatial representation rather than reasoning strategy. These findings reveal fundamental scalability limitations in current MLLMs and position TriViewBench as a controlled diagnostic framework for analyzing structural reasoning failures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.