2607.17790v1 Jul 20, 2026 cs.CV

ReViV: 단일 시점 근접 영상에서 관찰자와 환경을 4차원으로 재구성하는 방법

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

Marc Pollefeys
Marc Pollefeys
Citations: 705
h-index: 13
X. Lyu
X. Lyu
Citations: 0
h-index: 0
Zhiyin Qian
Zhiyin Qian
Citations: 417
h-index: 3
Gen Li
Gen Li
Citations: 144
h-index: 5
Xucong Zhang
Xucong Zhang
Citations: 4,925
h-index: 24
Siyu Tang
Siyu Tang
Citations: 192
h-index: 5

웨어러블 전면 카메라와 같은 근접 장치는 인간 관찰자와 주변 환경 간의 지속적인 상호작용을 포착하기 위한 독특한 시점을 제공합니다. 이러한 4차원 표현을 재구성할 수 있는 통합적이고 효율적인 다중 모드 모델은 매우 바람직합니다. 그러나 기존 접근 방식은 종종 사전 계산된 카메라 경로와 같은 추가 입력에 의존하거나, 장면 인식과 인간의 자기 운동 모델링을 별개의 문제로 취급하여 이들의 강한 상호 의존성을 간과하며, 느린 추론 속도를 갖는 경우가 많습니다. 이러한 한계를 해결하기 위해, 본 논문에서는 단일 모노 렌즈 RGB 영상에서 관찰자와 환경 모두의 동적 정보를 추출하는 최초의 통합된 근접 4차원 재구성 프레임워크인 ReViV를 제시합니다. 본 연구는 RGB 영상, 카메라 경로, 시선 방향, 전신 동작, 손 동작, 깊이 정보 등 다양한 다중 모드 신호에 대한 전체 결합 확률 분포를 학습하는 문제로 정의합니다. Masked Generative Egocentric Transformer를 활용하여 ReViV는 단일 순방향 아키텍처 내에서 작동하며, 빠른 추론 속도로 관찰자와 환경 모두에 대해 시간적으로 일관된 4차원 재구성을 동시에 수행합니다. HoloAssist, HOT3D, ARCTIC, Aria Digital Twin 및 TACO를 포함한 다양한 벤치마크에서의 광범위한 실험 결과는 ReViV가 전체적인 자아-신체, 손, 시선 재구성, 카메라 추적 분야에서 최첨단 정확도와 효율성을 달성하며, 특정 작업에 대한 과도한 사전 지식 없이 경쟁력 있는 근접 깊이 추정 성능을 유지함을 보여줍니다. 코드 및 모델은 완전 공개되어 있습니다: https://reviv4d.github.io/.

Original Abstract

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!