MSUE: 다중 모드 축구 이해 전문가
MSUE: Multi-Modal Soccer Understanding Expert
본 논문에서는 2026년 SoccerNet VQA Challenge에 대한 당사의 솔루션을 제시합니다. 먼저, 비전-언어 모델(VLM)을 기반으로 비용 효율적인 데이터 합성 파이프라인을 개발하여 원시 도메인 데이터를 체계적으로 재구성하고 다양한 VQA 샘플을 생성합니다. 이러한 샘플에는 간결한 답변과 장문의 응답이 포함됩니다. 둘째, 우리는 MSUE를 제안합니다. MSUE는 대규모 언어 모델(LLM)을 사용하여 질문을 텍스트, 이미지 및 비디오 전문가에게 동적으로 할당하는 다중 전문가 질의응답 아키텍처입니다. 이러한 전문가는 각각 강력한 텍스트 기반 모델인 Gemini3-Flash, 미세 조정된 Qwen3-VL, 그리고 외부 지식 베이스로 구현되어 협력적으로 VQA 성능을 향상시킵니다. MSUE는 Challenge 벤치마크에서 정확도 extbf{0.95}를 달성하여 리더보드에서 3위를 차지했습니다.
This paper presents our solution to the 2026 SoccerNet VQA Challenge. We first develop a cost-effective data synthesis pipeline driven by a Vision-Language Model (VLM), which systematically restructures raw domain data into diverse VQA samples, including concise answers and long-form responses. Second, we propose MSUE, a multi-expert question answering architecture that employs a Large Language Model (LLM) to dynamically dispatch questions to text, image, and video experts. These experts are instantiated as a strong text baseline Gemini3-Flash, a fine-tuned Qwen3-VL, and an external knowledge base, respectively, working collaboratively to enhance VQA performance. MSUE achieves an accuracy of \textbf{0.95} on the challenge benchmark, securing third place in the leaderboard.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.