2607.28751v1 Jul 30, 2026 cs.CV

ReLoop-UME: 학습 가능한 검색 레지스터를 갖춘 순환 심층 구조를 이용한 범용 다중 모드 임베딩

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

Xinyu Tang
Xinyu Tang
Citations: 4,823
h-index: 8
Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shijie Wang
Shijie Wang
Citations: 11,489
h-index: 5
Haiyun Guo
Haiyun Guo
Citations: 1,713
h-index: 20
Xiangzhao Hao
Xiangzhao Hao
Citations: 34
h-index: 4
Guangyun Cao
Guangyun Cao
Citations: 0
h-index: 0

범용 다중 모드 임베딩(UME)은 서로 다른 다중 모드 입력을 공유된 임베딩 공간으로 매핑합니다. 기존의 UME 모델들은 단일 순방향 인코딩을 통해 임베딩을 생성하거나, 명시적인 추론 토큰과 잠재적 자기 회귀 상태를 추가하여 계산을 수행합니다. 토큰 확장은 복잡한 매칭 성능을 향상시킬 수 있지만, 순차적인 생성이 검색 지연 시간을 증가시키고 최종 임베딩이 생성된 중간 상태에 의존하게 만듭니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 모델 깊이를 따라 유용한 계산을 확장하면서 토큰 공간을 고정할 수 있는지 질문합니다. 독립적으로 학습된 UME 모델의 각 레이어에서 양수-음수 유사성 분리를 분석한 결과, 일관된 패턴이 나타나는 것을 확인했습니다: 초기 레이어는 다중 모드 입력을 문맥화하고, 연속적인 중간-후반 레이어는 검색 및 분류에 유용한 특징을 형성하며, 마지막 레이어는 이를 임베딩 공간으로 매핑합니다. 이러한 분석 결과를 바탕으로, 본 연구에서는 초기 레이어를 한 번 실행하고, 파라미터 공유 검색 블록을 재사용하여 순환적으로 처리하며, 마지막 매핑 레이어를 최종 단계에서 적용하는 ReLoop-UME 모델을 제안합니다. 학습 가능한 검색 레지스터는 지속적인 검색 관련 상태를 제공하며, 루프를 통해 정보를 누적하고 교환하여 최종 레지스터가 임베딩 결과를 생성합니다. MMEB-V2 및 MRMR 데이터셋에 대한 실험 결과, ReLoop-UME는 다양한 백본 구조에서 검색 성능을 향상시키며, UME-R1보다 44.9배 빠르고 PLUME보다 1.5배 빠른 속도를 보였습니다.

Original Abstract

Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings through single forward encoding or add computation through explicit rationale tokens and latent autoregressive states. Although token expansion can improve complex matching, serial generation increases retrieval latency and makes the final embedding depend on generated intermediate states. This raises a different question: can useful computation be expanded along model depth while keeping the token workspace fixed? We analyze positive-negative similarity separation at every layer of independently trained UME models and observe a shared progression: early layers contextualize multimodal inputs, a contiguous middle-to-late stage forms retrieval-discriminative features, and the final layers map them into the embedding space. Based on this finding, we propose ReLoop-UME, which executes the early layers once, recurrently reuses a parameter-shared retrieval-forming block, and applies the final mapping layers after the last loop. Learnable Retrieval Registers provide persistent retrieval-specific states that accumulate and exchange evidence across loops, with the final register serving as the embedding readout. On MMEB-V2 and MRMR, ReLoop-UME consistently improves retrieval across different backbones while running 44.9x faster than UME-R1 and 1.5x faster than PLUME.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!