EgoVis 2026의 첫 번째 EgoCross 챌린지: 교차 도메인 개인 시점 비디오 질의 응답
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
EgoCross는 다중 모달 대규모 언어 모델이 일상적인 상황을 넘어 일반화할 수 있는지를 평가하기 위해 설계된 교차 도메인 개인 시점 비디오 질의 응답 벤치마크입니다. 첫 번째 EgoCross 챌린지는 CVPR 2026에서 개최된 세 번째 EgoVis 워크숍에서 진행되었으며, 모델은 수술, 산업 조립, 익스트림 스포츠 및 동물 관점에서 촬영된 1인칭 비디오를 사용하여 평가되었습니다. 각 테스트 예시는 개인 시점 비디오 클립, 질문, 그리고 정답 후보 4개로 구성되며, 모델은 이 중에서 올바른 답을 선택해야 합니다. 본 기술 보고서는 챌린지 과제, 벤치마크 리소스 및 두 가지 공식 Codabench 트랙에 대해 소개합니다. 소스 제한(Source-Limited) 트랙에서는 참가자들이 공식 기본 모델과 작은 지원 데이터 세트에만 사용하도록 제한하며, 오픈 소스(Open-Source) 트랙은 더 넓은 범위의 모델과 훈련 데이터를 사용할 수 있도록 허용하지만, 대상 도메인에 대한 훈련 데이터를 직접 생성하는 것을 금지합니다. 총 1,500건 이상의 제출물이 접수되었으며, 130명 이상의 참가자가 참여했습니다. 오픈 소스 트랙에는 19개 팀이, 소스 제한 트랙에는 38개 팀이 참여했습니다. 본 보고서에서는 공식 리더보드 결과를 제시하고, 양쪽 트랙에서 우승한 솔루션을 요약합니다. 본 보고서가 교차 도메인 개인 시점 비디오 이해를 발전시키는 데 유용한 기술 참고 자료가 되기를 바랍니다. 챌린지 데이터, 기본 구현 및 우승 팀이 공개한 코드를 포함한 모든 리소스는 공개적으로 제공됩니다.
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answers, from which the model must select the correct option. This technical report introduces the challenge task, benchmark resources, and two official Codabench tracks. The Source-Limited Track restricts participants to the official baseline model and a small support set, whereas the Open-Source Track permits broader choices of models and training data under rules that prohibit the manual construction of target-domain training data. In total, the challenge received more than 1,500 submissions from over 130 participants, with 19 teams participating in the Open-Source Track and 38 teams in the Source-Limited Track. We further present the official leaderboard results and summarize the winning solutions from both tracks. We hope that this report will serve as a useful technical reference for advancing cross-domain egocentric video understanding. All resources, including the challenge data, baseline implementation, and code released by the winning teams, are made publicly available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.