2607.14616v1 Jul 16, 2026 cs.AI

SportD: 대규모 시각-언어 모델이 물리적 전략을 수립할 수 있을까?

SportD: How do VLMs physically strategize?

Addison J. Wu
Addison J. Wu
Citations: 139
h-index: 3
Haotian Xia
Haotian Xia
Citations: 96
h-index: 6
Jasin Cekinmez
Jasin Cekinmez
Citations: 0
h-index: 0
Akshay Bharadhwaj
Akshay Bharadhwaj
Citations: 0
h-index: 0
Anay Putty
Anay Putty
Citations: 0
h-index: 0
Anirudh Ravishankar
Anirudh Ravishankar
Citations: 0
h-index: 0
Jaewoong Lee
Jaewoong Lee
Citations: 0
h-index: 0
Jinglin Xiao
Jinglin Xiao
Citations: 0
h-index: 0
Kyumin Andrew Shim
Kyumin Andrew Shim
Citations: 0
h-index: 0
Mishika Ahuja
Mishika Ahuja
Citations: 0
h-index: 0
Nisarga Patil
Nisarga Patil
Citations: 0
h-index: 0
Zhuohan Liu
Zhuohan Liu
Citations: 0
h-index: 0
Weining Shen
Weining Shen
Citations: 92
h-index: 6
Leon Liu
Leon Liu
Citations: 0
h-index: 0

시각-언어 모델은 시각적인 장면을 해석하는 능력이 점점 더 발전하고 있지만, 이러한 모델들이 정보를 활용하여 전략적으로 효과적인 결정을 내릴 수 있는지 여부는 불분명합니다. 본 연구에서는 축구라는 분야에서 이 질문을 조사합니다. 모델은 공이 있는 상황에서 발생하는 몇 초 전의 영상을 관찰하고, 슛을 할 것인지 특정 동료에게 패스를 할 것인지 선택해야 합니다. 기존의 시각 정보 이해 작업과는 달리, 축구는 모든 가능한 행동의 가치를 추정함으로써 결정을 정량적으로 평가할 수 있도록 합니다. 본 연구에서는 2022 FIFA 월드컵에서 추출한 478개의 공이 있는 상황에서의 결정들을 포함하는 벤치마크인 SportD를 소개합니다. 각 모델의 선택은 공격 팀이 골을 넣을 확률을 가장 높이는 행동을 추정하는 '소유 가치' 모델과 비교됩니다. 이를 통해 최적 행동 정확도와 비최적 결정으로 인해 발생하는 손실을 측정할 수 있습니다. 세 가지 최첨단 시각-언어 모델에 대한 실험 결과, 가장 성능이 좋은 모델은 31.4%의 경우에 가장 높은 가치를 가진 행동을 선택하는 반면, 프로 선수들은 38.9%의 경우에 동일한 행동을 선택하며, 모든 모델에서 훨씬 더 큰 후회를 발생시킵니다. 추가 분석 결과, 시각-언어 모델은 낮은 변동성과 낮은 보상을 갖는 행동을 체계적으로 선호하는 것으로 나타났습니다. 즉, 모델들은 슛을 할 빈도가 낮고, 최적 정책이나 실제 선수들보다 상당히 덜 발전된 패스를 선택합니다. 또한, 모델들은 비최적인 행동이라도 실제 선수들의 특정 행동을 통계적으로 유의미하게 모방하는 경향이 있는데, 이는 기존 플레이 패턴에 대한 부분적인 모방일 가능성을 시사하며, 반사실적 대안에 대한 일관된 평가가 이루어지는 것은 아닐 수 있습니다. SportD는 시각-언어 모델에서 물리적인 전략적 추론 능력을 측정하기 위한 가치 기반 테스트 환경을 제공합니다.

Original Abstract

Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1415 decision scenarios across professional men's and women's soccer games, where a VLM observes the seconds before a decision and chooses the next action. Models only select the optimal action around 30% of the time, even less frequently than humans do. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 83-92% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed (Spearman corr. 0.30 to +0.52), despite no such relationship in the ground truth (Spearman corr. -0.08). The conservatism therefore reflects a mis-calibration of value. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!