2606.19253v1 Jun 17, 2026 cs.CV

OneCanvas: 파노라마 재투영을 통한 3차원 장면 이해

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

Dave Zhenyu Chen
Dave Zhenyu Chen
Technical University of Munich
Citations: 1,606
h-index: 11
Bartlomiej Baranowski
Bartlomiej Baranowski
Citations: 56
h-index: 1
Matthias Nießner
Matthias Nießner
Citations: 45
h-index: 3

기존의 비전-언어 모델(VLMs)에서 3차원 장면 이해를 위한 접근 방식은 복잡하고 모델에 특화된 기하학적 인코더에 의존하거나, 공간 추론 능력을 향상시키기 위해 대규모의 학습 데이터를 필요로 합니다. OneCanvas는 모든 시점에서 추출한 패치 특징들을 단일 평면 직교(equirectangular) 파노라마 캔버스 위에 통합합니다. 구체적으로, 각 패치는 깊이 정보와 카메라 자세를 사용하여 3차원 월드 좌표로 투영되고, 캔버스의 원점을 기준으로 해당 지점의 연속적인 위도와 경도에 따라 캔버스 상에 배치됩니다. 이때, 라스터화 또는 중복되는 시각 간의 통합 과정은 생략됩니다. 패치의 미터법 좌표를 나타내는 3차원 위치 임베딩이 특징 벡터에 추가되어 월드 포지션을 각도 기반 캔버스 좌표로 변환하는 과정에서 손실될 수 있는 깊이 정보를 복구합니다. 결과적으로, 모든 프레임의 패치들이 별도의 융합 또는 백본 아키텍처의 주요 수정 없이 하나의 공간 좌표계를 공유하게 됩니다. 사전 학습된 VLM은 이 표현을 일반적인 이미지로 간주하여 사용합니다. 캔버스는 관심 지점의 어떤 자세를 기준으로 정렬할 수 있으므로, 동일한 표현이 특정 시점에서 이루어지는 상황 추론을 직접적으로 지원하며, 이는 로봇 공학 및 인공지능 분야에서 흔히 요구되는 기능입니다. 이러한 표현 덕분에 우리는 공간 사전 학습 커리큘럼을 도입할 수 있습니다: 실제 이미지에서 추출한 객체의 패치 특징들을 선택된 3차원 월드 위치에 배치하여, 그렇지 않으면 비어 있는 캔버스 위에 생성함으로써, 다양한 공간 추론 과제를 위한 즉석 감독 신호를 제공합니다. 이때, 정답 분포를 제어하여 공간 추론의 간편한 방법을 줄입니다. OneCanvas는 SQA3D 및 VSI-Bench에서 최고 수준의 정확도를 달성했으며, SPBench 데이터셋에서 일반화 성능을 보였습니다. 또한, 가장 강력한 경쟁 모델보다 10배 적은 학습 컴퓨팅 자원을 사용했습니다.

Original Abstract

Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto a single equirectangular panoramic canvas. Namely, each patch is unprojected to a 3D world coordinate using its depth and camera pose, then placed on the canvas at the continuous longitude and latitude of that point as seen from the canvas origin, with no rasterization or aggregation across overlapping views. A 3D position embedding of the patch's metric coordinates is added to its feature, restoring the depth lost when collapsing the world position to an angular canvas coordinate. Patches from all frames thus share one spatial coordinate system with no fusion or major architectural modifications of the backbone. The pretrained VLM consumes this representation as if it were an ordinary image. Because the canvas can be centered on any pose of interest, the same representation directly supports situated reasoning from a specific viewpoint, a common requirement in robotics and embodied AI. Thanks to this representation, we can also introduce a spatial pretraining curriculum: by procedurally placing patch features of objects, drawn from real images, at chosen 3D world positions on an otherwise empty canvas, we generate on-the-fly supervision spanning a broad range of spatial reasoning tasks, with answer distributions controlled to reduce spatial reasoning shortcuts. OneCanvas achieves state-of-the-art accuracy on SQA3D and VSI-Bench, and generalizes to out-of-distribution data on SPBench, using an order of magnitude less training compute than the strongest competing methods.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!