2608.05774v1 Aug 06, 2026 cs.CV

SR-JEPA: 3차원 장면에서 예측적 잠재 상태 학습

SR-JEPA: Learning Predictive Latent State in 3D Scenes

Qifu Wen
Qifu Wen
Citations: 2
h-index: 1
Zihan Zhou
Zihan Zhou
Citations: 80
h-index: 3
Xianzhu Zeng
Xianzhu Zeng
Citations: 1
h-index: 1

Joint-embedding predictive architectures(JEPAs)는 누락된 관찰의 잠재 표현을 예측함으로써 학습하지만, 많은 JEPAs는 생성하는 인코더를 중심으로 평가됩니다. 본 연구에서는 훈련된 예측 경로 자체가 원본 3차원 장면에서 전체 객체가 없을 때 무엇을 추론하는지 질문합니다. 우리는 SR-JEPA를 제안하는데, 이는 장면 수준의 포인트 클라우드에 대한 포인트 기반 JEPA로, 원래의 고정된 예측 경로를 특정 위치에서 쿼리할 수 있습니다. 평가 시, 하나의 객체의 모든 포인트를 제거하고 인코딩하기 전에 동일한 모양이 없는 32포인트 쿼리로 교체하며, 이 쿼리는 해당 포인트의 중심에 배치됩니다. 학습은 독립적인 3차원 EMA 목표만을 사용하며, 재구성, 의미 레이블, 언어 또는 확장된 2차원 특징을 사용하지 않습니다. 5,953개의 보류된 ARKitScenes 객체에서, 복구된 잠재 표현은 43.13%의 의미 일치 매크로 정확도를 달성했으며, 이는 최고 수준보다 22.18점 높은 수치입니다. 예측 경로를 무작위화하면 9.78점이 감소하고, 일치하는 도너 컨텍스트를 대체하면 21.98점이 감소합니다. 8,570개의 Sr3D 지원 쌍에서, 전체 잠재 표현은 41.15의 AP(Average Precision)를 달성했습니다. 누락된 객체의 잠재 표현에서 디코딩된 동일성 정보와 함께 기준 객체의 동일성 및 기하학 정보를 결합하면 39.37의 AP를 달성하며, 여전히 해결되지 않은 1.78점의 잔여 오차가 존재합니다. 이러한 결과는 쿼리 가능한 조립형 3차원 예측 상태를 보여줍니다. 즉, 모델은 컨텍스트에 따라 달라지는 객체 내용을 완성하며, 이는 하위 계산에서 기하학적 정보와 결합됩니다.

Original Abstract

Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. We introduce SR-JEPA, a point-native JEPA for scene-scale point clouds whose original frozen predictive pathway can be queried at a supplied location. At evaluation, every point of one object is removed before encoding and replaced by the same shape-free 32-point query at its centroid. Training uses only self-contained 3D EMA targets: no reconstruction, semantic labels, language, or lifted 2D features. On 5,953 held-out ARKitScenes objects, the imputed latent reaches 43.13% semantic-identity macro accuracy, 22.18 points above the strongest floor. Randomizing the prediction path removes 9.78 points, while substituting matched donor context removes 21.98 points. On 8,570 Sr3D support pairs, the full latent reaches 41.15 AP; identity decoded from the missing-object latent, combined with anchor identity and geometry, reaches 39.37 AP, leaving an unresolved 1.78-point residual. These results reveal a queryable, compositional 3D predictive state: the model completes context-dependent entity content, which downstream computation combines with metric geometry.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!