2604.10391v1 Apr 12, 2026 cs.CV

FishRoPE: 전방향 시각 인식을 위한 투영 회전 위치 임베딩

FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception

S. Yogamani
S. Yogamani
Citations: 7,784
h-index: 41
Rahul Ahuja
Rahul Ahuja
Citations: 59
h-index: 3
Mudit Jain
Mudit Jain
Citations: 73
h-index: 4
B. Sudhakar
B. Sudhakar
Citations: 24
h-index: 3
Venkatraman Narayanan
Venkatraman Narayanan
Citations: 17
h-index: 2
Pratik Likhar
Pratik Likhar
Citations: 34
h-index: 2
Varun Ravi Kumar
Varun Ravi Kumar
Citations: 42
h-index: 4

비전 기반 모델(VFMs)과 버드아이뷰(BEV) 표현 방식은 시각 인지 분야에 큰 발전을 가져왔지만, 이들의 내부 공간 표현은 핀홀 카메라의 직교 기하 구조를 가정합니다. 반면, 주변 시야 확보를 위해 자율 주행 차량에 널리 사용되는 어안 렌즈는 심각한 방사 왜곡을 나타내어 이러한 표현 방식을 기하학적으로 일관성 없게 만듭니다. 또한, 대규모 어안 렌즈 데이터셋의 부족으로 인해 기존 모델을 처음부터 재학습하는 것은 비현실적입니다. 본 논문에서는 두 가지 구성 요소로 구성된 경량 프레임워크인 exttt{FishRoPE}를 제안합니다. 첫 번째는 LoRA(Low-Rank Adaptation)를 사용하여 사전 학습된 DINOv2 모델의 풍부한 자기 지도 학습 특징을 어안 렌즈 데이터에 적용하며, 특정 작업에 대한 사전 학습 없이도 가능합니다. 두 번째는 어안 렌즈 투영의 구면 좌표계를 사용하여 어텐션 메커니즘을 재파라미터화하는 Fisheye Rotary Position Embedding(FishRoPE)으로, 자기 어텐션 및 크로스 어텐션이 픽셀 거리 대신 각도 차이를 기반으로 작동하도록 합니다. FishRoPE는 아키텍처에 구애받지 않으며, 계산 오버헤드가 미미하고, 핀홀 기하 구조에서는 표준 방식으로 작동합니다. 저희는 FishRoPE를 WoodScape 2D 객체 탐지(54.3 mAP) 및 SynWoodScapes BEV 분할(65.1 mIoU)에 적용하여, 두 가지 벤치마크에서 최고 성능을 달성했습니다.

Original Abstract

Vision foundation models (VFMs) and Bird's Eye View (BEV) representation have advanced visual perception substantially, yet their internal spatial representations assume the rectilinear geometry of pinhole cameras. Fisheye cameras, widely deployed on production autonomous vehicles for their surround-view coverage, exhibit severe radial distortion that renders these representations geometrically inconsistent. At the same time, the scarcity of large-scale fisheye annotations makes retraining foundation models from scratch impractical. We present \ours, a lightweight framework that adapts frozen VFMs to fisheye geometry through two components: a frozen DINOv2 backbone with Low-Rank Adaptation (LoRA) that transfers rich self-supervised features to fisheye without task-specific pretraining, and Fisheye Rotary Position Embedding (FishRoPE), which reparameterizes the attention mechanism in the spherical coordinates of the fisheye projection so that both self-attention and cross-attention operate on angular separation rather than pixel distance. FishRoPE is architecture-agnostic, introduces negligible computational overhead, and naturally reduces to the standard formulation under pinhole geometry. We evaluate \ours on WoodScape 2D detection (54.3 mAP) and SynWoodScapes BEV segmentation (65.1 mIoU), where it achieves state-of-the-art results on both benchmarks.

3 Citations
1 Influential
20.5 Altmetric
107.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!