2608.04515v1 Aug 05, 2026 cs.CV

CARVE: 효율적인 3차원 의료 영상 분석을 위한 시각적 증거의 횡단면 등방성 재분배

CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding

Zhenyu Yi
Zhenyu Yi
Citations: 25
h-index: 2
Qiang Hu
Qiang Hu
Citations: 33
h-index: 3
Yusong Sun
Yusong Sun
Citations: 107
h-index: 4
Zhenhao Li
Zhenhao Li
Citations: 0
h-index: 0
Jiaxuan Zhao
Jiaxuan Zhao
Citations: 0
h-index: 0
Lichi Zhang
Lichi Zhang
Citations: 97
h-index: 6

슬라이스 기반 MLLM(대규모 언어 모델)은 성숙한 2D 인코더를 활용하여 3차원 볼륨을 2D 슬라이스의 시퀀스로 표현합니다. 그러나 이러한 슬라이스 단위의 방식은 수천 개의 시각적 토큰을 생성하며, 이는 LLM의 성능에 부담을 주고, 많은 토큰이 인접한 슬라이스 간의 중복된 시각 정보를 포함합니다. 우리는 3차원 의료 VQA(시각 질의응답) 벤치마크 두 가지를 사용하여 증가하는 시각적 토큰 예산이 성능 향상에 미치는 영향을 분석하고, 그 결과로 얻어지는 이점은 감소하는 경향을 보임을 확인했습니다. 즉, 비용은 계속 상승하지만 정확도는 포화 상태에 도달하며, 동일한 예산 내에서 더 높은 인체 내 해상도를 확보하는 것이 슬라이스 수를 늘리는 것보다 효과적입니다. 따라서 예산을 단순히 확장하기보다는 보다 선택적으로 할당해야 합니다. 그러나 대부분의 토큰 압축 방법은 주로 2D 이미지 또는 비디오에 적용되는데, 이는 공간적인 배열이나 시간적인 움직임에서 비롯되는 중복성을 처리하는 데 초점을 맞추고 있습니다. 본 논문에서는 CARVE라는 학습이 필요 없는 프레임워크를 제시합니다. 이 프레임워크는 LLM 추론 전에 시각적 토큰을 압축하고, 토큰 감소를 예산 제약 하에 이루어지는 2.5차원 할당 문제로 재정의합니다. CARVE는 깊이 축을 일관성 있는 창으로 분할하고, 정규화된 슬라이스 간 증거량에 따라 토큰을 비균등하게 할당합니다. 공유된 예산 하에서 CARVE는 대표적인 슬라이스에 공간적 앵커를 구축하고 전체 볼륨에서 지역적으로 다양한 증거를 검색한 다음, 각 창 내의 인접한 앵커에 나머지 유효한 토큰을 병합합니다. Hulu-Med-7B 데이터셋에서 약 80%의 시각적 토큰을 제거하면서 CARVE는 모든 압축 기준 모델보다 우수한 성능을 보이며, 가장 강력한 기준 모델 대비 6.2점 높은 품질 유지율을 달성하고, 세 가지 VQA 벤치마크에서 전체 토큰 성능의 98.1%를 유지합니다.

Original Abstract

Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture overlapping visual evidence across adjacent slices. To understand how effectively a growing visual token budget improves performance, we perform scaling analyses on two 3D medical VQA benchmarks and find diminishing returns: cost keeps rising while accuracy saturates, and improving in-plane resolution is more effective than adding slices at comparable budgets. The budget should therefore be allocated more selectively rather than simply enlarged, yet most token compression methods are designed for 2D images or videos, where redundancy arises from spatial layout or temporal motion rather than from near-duplicate content along the depth axis. We present CARVE, a training-free framework that compresses visual tokens prior to LLM inference and casts token reduction as budget-constrained 2.5D allocation. CARVE partitions the depth axis into coherent windows and allocates tokens non-uniformly according to normalized cross-slice evidence. Under a shared budget, CARVE builds spatial anchors on representative slices and retrieves locally varying evidence from the full volume, then merges remaining eligible tokens into nearby anchors within each window. Removing roughly 80% of the visual tokens on Hulu-Med-7B, CARVE leads all compression baselines on every AMOS-MM report-generation metric, with 6.2 points higher retention of full-token quality than the strongest baseline, and preserves 98.1% of full-token performance across three VQA benchmarks.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!