2607.29033v1 Jul 31, 2026 cs.CV

SAM+D: 깊이 기반 경로 LoRA 및 깊이 이동을 통한 매개변수 효율적인 SAM 계열 모델의 차원 확장

SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

Hao Sun
Hao Sun
Citations: 26
h-index: 2
Shiyu Teng
Shiyu Teng
Citations: 291
h-index: 7
Yu Song
Yu Song
Citations: 34
h-index: 3
Yen-wei Chen
Yen-wei Chen
Citations: 304
h-index: 10
I. Nishikawa
I. Nishikawa
Citations: 961
h-index: 14

기존의 SAM과 같은 2차원 기초 모델을 3차원으로 확장하는 방법은 슬라이스를 독립적으로 처리하여 슬라이스 간의 문맥 정보를 무시하거나, 상당한 아키텍처 변경 및 재학습이 필요합니다. 본 논문에서는 extbf{SAM+D}라는 매개변수 효율적인 프레임워크를 제시합니다. SAM+D는 2차원 SAM 모델을 공간적으로 한 차원 확장하여 3차원 볼륨 분할을 가능하게 하며, 또한 매개변수 효율적인 미세 조정 방식을 통해 비디오 기반 SAM2에서 처음으로 4차원(3차원 + 시간) 시공간 분할을 수행합니다. 이 과정에서 대부분의 사전 학습된 파라미터는 고정됩니다. SAM+D는 고정된 트랜스포머 블록에 두 개의 가벼운, 모델 독립적인 모듈을 도입합니다: (1) 공간적으로 적응 가능한 저랭크 업데이트를 위한 학습 기반 경로를 가진 extbf{깊이 기반 경로 LoRA (DRLoRA)} 전문가 모듈과 (2) 추가 파라미터 비용 없이 슬라이스 간 특징 교환을 수행하는 extbf{깊이 이동 모듈 (DSM)}. 이 두 가지 모듈은 함께 볼륨 수준의 문맥 정보를 제공하면서 SAM의 경우 약 2.8%, SAM2의 경우 약 3.7%의 파라미터만 미세 조정합니다. 우리는 SAM+D를 두 가지 다른 환경에서 평가했습니다. 첫째, 3차원 분할 작업에서는 SAM(2D → 3D) 모델을 KiTS, Pancreas, LiTS 및 Colon이라는 네 가지 CT 데이터셋으로 평가하고, 둘째, 4차원 분할 작업에서는 SAM2 (2D+T → 3D+T) 모델을 세포 추적 챌린지(CTC) 데이터셋(Fluo-N3DH-SIM+)으로 평가했습니다. 두 환경 모두에서 SAM+D는 단일 프롬프트 설정에서 기존 방법보다 경쟁력 있거나 더 나은 결과를 달성했으며, 더 적은 수의 학습 가능한 파라미터를 사용합니다. 이러한 결과는 SAM+D가 SAM 계열 아키텍처, 대상 차원(3차원, 4차원) 및 의료 영상에서 생체 장면 이해에 이르는 다양한 도메인 전반에 걸쳐 일반화될 수 있음을 보여줍니다. 코드 공개 저장소: https://github.com/JerrySongCST/SAM-Plus-D.

Original Abstract

Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or require substantial architectural changes and retraining. In this paper, we present \textbf{SAM+D}, a parameter-efficient framework that lifts SAM-family models by one spatial dimension---enabling 3D volumetric segmentation from 2D SAM and, for the first time via parameter-efficient fine-tuning, end-to-end 4D (3D+T) spatiotemporal segmentation from video-based SAM2---while keeping the vast majority of pre-trained parameters frozen. SAM+D introduces two lightweight, model-agnostic modules into frozen transformer blocks: (1)~\textbf{Depth-Routed LoRA (DRLoRA)} experts with learned routing for spatially adaptive low-rank updates, and (2)~\textbf{Depth Shift Modules (DSM)} for cross-slice feature exchange at zero additional parameter cost. Together, they provide volume-level context while tuning only ${\sim}$2.8\% of parameters for SAM and ${\sim}$3.7\% for SAM2. We evaluate SAM+D in two distinct settings, each lifting the base model by one spatial dimension: 3D segmentation, where SAM(2D$\,\to\,$3D) is evaluated on four CT benchmarks (KiTS, Pancreas, LiTS, Colon), and 4D segmentation, where SAM2 (2D+T$\,\to\,$3D+T) is evaluated on a cell tracking challenge (CTC) dataset (Fluo-N3DH-SIM+). In both settings SAM+D achieves competitive or superior results under the single-point prompt setting while using fewer trainable parameters than existing methods, demonstrating that SAM+D generalizes across SAM-family architectures, target dimensionalities (3D, 4D), and domains spanning medical imaging and bio-scene understanding. Code is publicly available at https://github.com/JerrySongCST/SAM-Plus-D.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!