보는 것이 곧 결정하는 것인가? 다중 모드 LLM이 효과적인 CEO 역할을 수행할 수 있는가?
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
최근 대규모 언어 모델(LLM)은 자율적 의사 결정 에이전트로 점점 더 많이 활용되고 있습니다. 그러나 경영 활동에서의 의사 결정에는 기존 벤치마크가 주로 텍스트 기반 환경에 국한되어 있어, 모델이 시각적인 비즈니스 정보를 인식하고 이를 효과적으로 통합하여 의사 결정의 질을 향상시킬 수 있는지 불분명합니다. 본 연구에서는 C-SUITEBENCH라는 통제된 다중 모드 벤치마크를 소개합니다. 이 벤치마크는 50개의 시나리오에서 텍스트 기반 조건과 다중 모드 조건을 함께 제공하는 5가지 의사 결정 과제를 포함하고 있습니다. 우리는 9개의 최첨단 모델을 CEO 역할을 수행하도록 배치하고, 그들의 의사 결정 능력을 평가했습니다. 다중 모드 입력은 일관되게 증거 중심적인 추론을 향상시키며, 특히 위험 예측 및 이사회 보고와 같은 영역에서 가장 큰 개선 효과를 보였습니다. 그러나 우리는 다중 모드 통합의 역설을 발견했습니다: 시각적 비즈니스 정보를 추가하면 모든 9개의 모델에서 제한된 자원 할당이 저하됩니다. 이는 시각적 정보 자체가 의사 결정에 도움이 되더라도, 제약 조건 만족도를 낮추는 결과를 초래합니다. 분석 결과, 이러한 실패는 각 시각 채널이 개별적으로 도움을 주지만, 이들의 조합은 디코딩 과정에서 제약 조건을 위반하는 '정보 과부하' 현상으로 인해 발생함을 확인했습니다. 본 연구의 결과는 다중 모드 에이전트에서 시각적 인식과 제약 조건 기반 행동이 분리된 병목 지점이며, 무분별한 시각 정보 추가가 중요한 의사 결정에 해를 끼칠 수 있음을 보여줍니다. 이러한 결과를 바탕으로 향후 경영 AI 시스템을 위한 선택적인 시각 정보 활용 전략의 필요성을 강조합니다.
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.