2608.04589v1 Aug 05, 2026 cs.CV

EgoVis 2026의 첫 번째 EgoCross 챌린지: 교차 도메인 개인 시점 비디오 질의 응답

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Liqiang Nie
Liqiang Nie
Citations: 143
h-index: 7
Kunyu Peng
Kunyu Peng
Citations: 1,753
h-index: 21
Leyi Wu
Leyi Wu
Citations: 49
h-index: 3
D. Paudel
D. Paudel
Citations: 4,035
h-index: 32
L. V. Gool
L. V. Gool
Citations: 4,766
h-index: 31
Licheng Jiao
Licheng Jiao
Citations: 946
h-index: 16
Yongqin Xian
Yongqin Xian
Citations: 3,536
h-index: 10
Zhiheng Fu
Zhiheng Fu
Citations: 519
h-index: 13
Zixu Li
Zixu Li
Citations: 567
h-index: 14
A. Tonioni
A. Tonioni
Citations: 4,969
h-index: 23
Tianwen Qian
Tianwen Qian
Citations: 557
h-index: 9
Yupeng Hu
Yupeng Hu
Citations: 735
h-index: 15
Yu Li
Yu Li
Citations: 353
h-index: 4
Xu Zheng
Xu Zheng
Citations: 46
h-index: 4
Xiaoling Wang
Xiaoling Wang
Citations: 38
h-index: 3
Federico Tombari
Federico Tombari
Citations: 1,504
h-index: 18
Ying-Cong Chen
Ying-Cong Chen
Citations: 172
h-index: 2
Wenbo Wang
Wenbo Wang
Citations: 9
h-index: 2
Taku Murakawa
Taku Murakawa
Citations: 0
h-index: 0
Toru Tamaki
Toru Tamaki
Citations: 13
h-index: 2
Yi Wen
Yi Wen
Citations: 10
h-index: 3
Zhenglin Du
Zhenglin Du
Citations: 49
h-index: 4
Zhengyang Li
Zhengyang Li
Citations: 4
h-index: 1

EgoCross는 다중 모달 대규모 언어 모델이 일상적인 상황을 넘어 일반화할 수 있는지를 평가하기 위해 설계된 교차 도메인 개인 시점 비디오 질의 응답 벤치마크입니다. 첫 번째 EgoCross 챌린지는 CVPR 2026에서 개최된 세 번째 EgoVis 워크숍에서 진행되었으며, 모델은 수술, 산업 조립, 익스트림 스포츠 및 동물 관점에서 촬영된 1인칭 비디오를 사용하여 평가되었습니다. 각 테스트 예시는 개인 시점 비디오 클립, 질문, 그리고 정답 후보 4개로 구성되며, 모델은 이 중에서 올바른 답을 선택해야 합니다. 본 기술 보고서는 챌린지 과제, 벤치마크 리소스 및 두 가지 공식 Codabench 트랙에 대해 소개합니다. 소스 제한(Source-Limited) 트랙에서는 참가자들이 공식 기본 모델과 작은 지원 데이터 세트에만 사용하도록 제한하며, 오픈 소스(Open-Source) 트랙은 더 넓은 범위의 모델과 훈련 데이터를 사용할 수 있도록 허용하지만, 대상 도메인에 대한 훈련 데이터를 직접 생성하는 것을 금지합니다. 총 1,500건 이상의 제출물이 접수되었으며, 130명 이상의 참가자가 참여했습니다. 오픈 소스 트랙에는 19개 팀이, 소스 제한 트랙에는 38개 팀이 참여했습니다. 본 보고서에서는 공식 리더보드 결과를 제시하고, 양쪽 트랙에서 우승한 솔루션을 요약합니다. 본 보고서가 교차 도메인 개인 시점 비디오 이해를 발전시키는 데 유용한 기술 참고 자료가 되기를 바랍니다. 챌린지 데이터, 기본 구현 및 우승 팀이 공개한 코드를 포함한 모든 리소스는 공개적으로 제공됩니다.

Original Abstract

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answers, from which the model must select the correct option. This technical report introduces the challenge task, benchmark resources, and two official Codabench tracks. The Source-Limited Track restricts participants to the official baseline model and a small support set, whereas the Open-Source Track permits broader choices of models and training data under rules that prohibit the manual construction of target-domain training data. In total, the challenge received more than 1,500 submissions from over 130 participants, with 19 teams participating in the Open-Source Track and 38 teams in the Source-Limited Track. We further present the official leaderboard results and summarize the winning solutions from both tracks. We hope that this report will serve as a useful technical reference for advancing cross-domain egocentric video understanding. All resources, including the challenge data, baseline implementation, and code released by the winning teams, are made publicly available.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!