2607.18603v1 Jul 21, 2026 cs.IR

AutoIndex: Learning Representation Programs for Retrieval

Andrew Drozdov
Andrew Drozdov
University of Massachusetts Amherst
Citations: 767
h-index: 13
Sam O’Nuallain
Sam O’Nuallain
Citations: 0
h-index: 0
Nithya Rajkumar
Nithya Rajkumar
Citations: 0
h-index: 0
Ramya Narayanasamy
Ramya Narayanasamy
Citations: 0
h-index: 0
Hanna Jiang
Hanna Jiang
Citations: 0
h-index: 0
Shreyas Chaudhari
Shreyas Chaudhari
Citations: 25
h-index: 2

We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the representations exposed to a retrieval system. Rather than tuning retrievers, rerankers, or a small set of preprocessing hyperparameters, AutoIndex searches over programs that slice, enrich, normalize, reweight, or reorganize documents before indexing. At each iteration, AutoIndex performs validation-guided program search, in which agents diagnose failures of the current program and synthesize candidate updates, retaining only updates that improve retrieval quality under the resulting index. We evaluate AutoIndex on CRUMB, a benchmark of heterogeneous retrieval tasks, with BM25 held fixed across all experiments. The learned programs improve recall over a static full-document BM25 baseline on all 8 tasks, with average gains of +8.4% in Recall@100 and +8.3% in nDCG@10, and largest gains of +30.5% in Recall@100 and +43.6% in nDCG@10. These results suggest that document representation should not be treated as a fixed preprocessing choice made before retrieval begins, but as an explicit optimization target. Code to reproduce our results is available at https://github.com/auto-index/autoindex.

0 Citations
0 Influential
38.92453324894 Altmetric
194.6 Score
Original PDF
11

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!