Vinci2: 지속적인 개체 중심 영상에서 능동적인 지원 제공
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
인공지능 어시스턴트는 언제 사용자의 요청 없이 먼저 정보를 제공해야 할까요? 지속적인 개체 중심 영상은 풍부하고 변화하는 맥락을 제공하며, 이를 통해 단순히 반응하는 것이 아니라 능동적으로 지원하는 새로운 형태의 가능성을 제시합니다. 그러나 기존 방식들은 대부분 사용자 쿼리를 기다리거나, 감지된 모든 이벤트를 응답이 필요한 것으로 취급하여 사용자의 과거 행동, 현재 활동 또는 실제로 지원이 필요한 상황인지 고려하지 않습니다. 우리는 능동적인 지원을 맥락에 의존하는 의사 결정 문제로 재정의합니다. 즉, 에이전트는 발생하는 현상을 인식할 뿐만 아니라, 축적된 시간적 맥락을 분석하여 언제 그리고 어떻게 개입해야 할지 판단해야 합니다. 이를 위해, 우리는 기존의 반응형 어시스턴트 Vinci를 발전시켜 능동적인 지원 기능을 제공하는 시스템인 Vinci2를 제안합니다. 평가 측면에서, 우리는 지속적인 개체 중심 영상에서의 능동적인 지원을 위한 최초의 대규모 벤치마크인 EgoServe를 소개합니다. EgoServe는 즉각적인 안전 경고부터 장기적인 습관 코칭까지 4가지 시간적 기억 범위를 포괄하며, 10개의 서비스 카테고리로 구성된 3,000개 이상의 서비스 인스턴스를 포함합니다. 모델링 측면에서, 우리는 학습이 필요 없는 메모리 기반 에이전트인 EgoMemo를 제안합니다. EgoMemo는 다중 스케일의 시간 요약, 의미론적 지식 그래프 및 시각적 임베딩 아카이브라는 세 가지 상호 보완적인 기억 표현을 유지하며, 각 타임스텝마다 검색 증강 추론을 통해 지원이 필요한지 판단하고, 필요에 따라 맥락에 맞는 응답을 생성합니다. 실험 결과는 EgoMemo가 EgoServe에서 강력한 기준 성능을 보여주며 기존의 개체 중심 벤치마크에서도 경쟁력 있는 성능을 유지한다는 것을 입증합니다. 저희의 벤치마크 데이터와 코드는 https://sitonggong.github.io/EgoServe-page/ 에서 공개적으로 이용할 수 있습니다.
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at \href{https://sitonggong.github.io/EgoServe-page/}{Vinci2}.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.