TennisVAR: 스트로크 증거 기반의 다중 모드 대규모 언어 모델 - 테니스 동영상에서의 전술적 추론
TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos
스포츠 영상 이해는 이벤트 인식에서 벗어나, 개별 동작들이 경기 진행에 어떻게 영향을 미치는지 설명하는 방향으로 발전하고 있습니다. 하지만 기존의 테니스 동영상 분석 방법들은 개별 스트로크를 모델링하지 않거나, 전술적 의존성을 고려하지 않고 고수준 분석만 제공하거나, 또는 근본적인 이벤트에 대한 연결 없이 분석을 생성합니다. 이러한 인식-이해 간의 격차를 해소하기 위해, 우리는 스트로크 증거 기반의 전술적 추론이라는 새로운 랠리 레벨 작업을 제안합니다. 이 작업은 모델이 개방형 답변, 계층적 전술 레이블, 정렬된 지원 스트로크 시퀀스 및 결정적인 핵심 동작을 동시에 예측하도록 요구하며, 각 증거 스트로크는 라켓과 공의 접촉 프레임에 연결됩니다. 또한 우리는 11,189개의 랠리 동영상, 41,485개의 스트로크 이벤트, 25,429개의 전술적 단위 및 11,189개의 질문-답변 쌍을 포함하는 대규모의 전문가 주석 데이터셋인 TRACE (Tactical Reasoning with Action-Chain Evidence in Tennis)를 소개합니다. TRACE는 세부적인 스트로크 속성, 스트로크 간의 전술적 관계, 계층적 전술 어노테이션 및 증거 기반 질문을 통합하여 사실 인식, 전술적 이해 및 의사 결정 추론을 가능하게 합니다. TRACE를 기반으로, 우리는 TennisVAR (Tennis Video Action-chain Reasoner)라는 증거 기반의 다중 모드 대규모 언어 모델을 제안합니다. TennisVAR는 "이벤트-관계-증거-전술"이라는 추론 패러다임을 따르며, 이벤트 파싱 모듈은 연속적인 랠리를 명시적인 스트로크 이벤트 시퀀스로 변환하고, 전술적 그래프 기반의 시간 추론기는 랠리 진행 상황과 동일 선수 간의 의사 결정 의존성을 동시에 모델링하여 질문과 관련된 증거 및 결정적인 동작을 식별합니다.
Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling their tactical dependencies or generate high-level analyses without grounding them in the underlying events. To bridge this perception-to-understanding gap, we formulate stroke-evidence-grounded tactical reasoning, a new rally-level task that requires models to jointly predict an open-ended answer, a hierarchical tactic label, an ordered sequence of supporting strokes, and decisive key actions, with each evidence stroke anchored to its racket-ball contact frame. We further introduce TRACE (Tactical Reasoning with Action-Chain Evidence in Tennis), a large-scale expert-annotated benchmark containing 11,189 rally videos, 41,485 stroke events, 25,429 tactical units, and 11,189 question-answer pairs, which unifies fine-grained stroke attributes, cross-stroke tactical relations, hierarchical tactic annotations, and evidence-grounded questions across factual perception, tactical understanding, and decision reasoning. Building on TRACE, we propose TennisVAR (Tennis Video Action-chain Reasoner), an evidence-grounded multimodal large language model that follows an "event-relation-evidence-tactic" reasoning paradigm, where an Event Parsing Module converts continuous rallies into explicit stroke-event sequences while a Tactical Graph-Guided Temporal Reasoner jointly models rally progression and same-player decision dependencies to identify question-relevant evidence and decisive actions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.