2605.29507v1 May 28, 2026 cs.AI

Xetrieval: 밀집 검색의 작동 원리에 대한 메커니즘 기반 설명

Xetrieval: Mechanistically Explaining Dense Retrieval

Jiaqi Li
Jiaqi Li
Citations: 360
h-index: 7
Zhixi Cai
Zhixi Cai
Citations: 87
h-index: 6
Zilong Zheng
Zilong Zheng
Citations: 420
h-index: 8
Jun Bai
Jun Bai
Beijing Institute for General Artificial Intelligence
Citations: 91
h-index: 5
Yang Liu
Yang Liu
Citations: 11
h-index: 3
Yichi Zhang
Yichi Zhang
Citations: 4
h-index: 1
Taichuan Li
Taichuan Li
Citations: 11
h-index: 1
Zhuofan Chen
Zhuofan Chen
Citations: 25
h-index: 2
Zixiao Jia
Zixiao Jia
Citations: 63
h-index: 1
Wenge Rong
Wenge Rong
Citations: 72
h-index: 3

밀집 검색 시스템이 높은 관련성 점수를 부여하는 이유를 설명하는 것은, 검색 결정이 불투명한 고차원 임베딩을 통해 이루어지기 때문에 어렵습니다. 기존의 설명 방식은 주로 어휘 일치, 토큰 정렬 또는 사후적인 텍스트 근거와 같은 표면적인 신호에 초점을 맞추며, 따라서 임베딩 수준에서 밀집 검색 행동을 형성하는 잠재적인 요인에 대한 제한적인 통찰력을 제공합니다. 본 논문에서는 밀집 검색을 설명하기 위한 임베딩 수준의 메커니즘 기반 프레임워크인 Xetrieval을 제안합니다. Xetrieval은 먼저 경량의 추론 모듈을 도입하여 단일 순방향 패스만으로 임베딩 공간에서 Chain-of-Thought 추론을 직접적으로 근사화하며, 문장 임베딩에 추론 관련 정보를 추가하는 동시에 비용이 많이 드는 자기 회귀 생성을 피합니다. 그런 다음 이 추론 강화된 임베딩을 희소하고 사람이 이해할 수 있는 특징으로 분해하며, 각 특징은 일관된 자연어 설명을 갖습니다. Xetrieval은 다양한 문서 관점에서의 희소 특징 중복을 집계하여 개별 검색 결정에 대한 특징 수준의 설명을 제공합니다. 다양한 검색 시스템 및 벤치마크에 대한 실험 결과, Xetrieval은 일관되고 해석 가능한 특징을 발견하고, 더 강력한 쌍 수준의 개입 효과를 나타내며, 작업 수준의 특징 제어를 지원한다는 것을 보여줍니다. 프로젝트 페이지 및 소스 코드는 https://hihiczx.github.io/Xetrieval 에서 확인할 수 있습니다.

Original Abstract

Explaining why dense retrievers assign high relevance scores remains challenging because retrieval decisions are made through opaque high-dimensional embeddings. Existing explanations often focus on surface signals, such as lexical matches, token alignments, or post-hoc textual rationales, and thus provide limited insight into the latent factors that shape dense retrieval behavior at the embedding level. We propose \textit{Xetrieval}, an embedding-level mechanistic framework for explaining dense retrieval. \textit{Xetrieval} first introduces a lightweight reasoning internalizer that approximates Chain-of-Thought reasoning directly in the embedding space with a single forward pass, enriching sentence embeddings with reasoning-oriented information while avoiding expensive autoregressive generation. It then decomposes these reasoning-enhanced embeddings into sparse, human-interpretable features, each associated with a coherent natural language description. By aggregating sparse feature overlaps across multiple document-side views, \textit{Xetrieval} provides feature-level explanations of individual retrieval decisions. Experiments on diverse retrievers and benchmarks show that \textit{Xetrieval} uncovers coherent interpretable features, yields stronger pair-level intervention effects, and supports task-level feature steering. The project page and source code are available at https://hihiczx.github.io/Xetrieval .

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!