2606.26899v1 Jun 25, 2026 cs.AI

디퓨전 트랜스포머를 이용한 생성적 검색: 메트릭 기반 순서 학습 및 하이브리드 정책 선호도 최적화

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

Kun Xu
Kun Xu
Citations: 55
h-index: 4
Zhenwei An
Zhenwei An
Citations: 26
h-index: 3
Zhongtao Jiang
Zhongtao Jiang
Citations: 221
h-index: 6
Jiachen Zhang
Jiachen Zhang
Citations: 18
h-index: 2
Zhaojun Wang
Zhaojun Wang
Citations: 12
h-index: 2
Chenghao Liu
Chenghao Liu
Citations: 56
h-index: 3
Yu Zhang
Yu Zhang
Citations: 1,025
h-index: 1
Renzhi Wang
Renzhi Wang
Citations: 52
h-index: 5
Yuxiao Zhang
Yuxiao Zhang
Citations: 0
h-index: 0
Songfang Huang
Songfang Huang
Citations: 142
h-index: 3

임베딩 기반 검색은 공유 벡터 공간에서 쿼리와 항목 간의 유사성을 기준으로 항목을 순위화하며, 일반적으로 최고 점수를 가진 항목들을 반환하는 것을 목표로 합니다. 그러나 많은 실제 환경에서는 이러한 방식이 적합하지 않습니다. 특정 미세한 패턴을 나타내는 초기 집합(seed set)이 주어졌을 때, 대상 속성을 만족하면서 동시에 해당 패턴 내에 머무르는 추가적인 항목들이 필요합니다. 본 연구는 이를 '패턴 보존 속성 검색'으로 정의하고, 두 목표 사이의 상충 관계를 해결하고자 합니다. 초기 집합 평균화는 패턴을 유지하지만 낮은 속성 영역에 머물게 되는 반면, 전역 속성 검색은 관련 없는 패턴으로 이어질 수 있습니다. 우리는 연속적인 생성적 검색 방법을 사용하여 문제를 접근합니다. 모델은 항목 임베딩 시퀀스를 읽고, 최근접 이웃 탐색을 위한 쿼리 임베딩을 생성합니다. 본 연구에서는 MO-DiT+HPPO라는 단계별 프레임워크를 제안합니다. 이는 원시 시퀀스 사전 학습, 다중 도메인 메트릭 기반 순서 지속적 사전 학습, 중심점 미세 조정 및 HPPO(Hybrid-Policy Preference Optimization)로 구성됩니다. 메트릭 기반 순서 학습은 희소한 온라인 검색 레이블을 사용하여 패턴 내의 궤적을 생성하며, 이를 통해 모델이 다양한 도메인에서 메트릭 개선 방향을 학습하도록 합니다. HPPO는 온라인 교차 지표를 사용한 하이브리드 후보 풀 레이블링과 참조 기반 선호도 최적화를 통해 생성된 쿼리 분포를 실제 온라인 목표와 일치시킵니다. 파레토 쌍 필터는 동일 패턴의 순수성을 저하시키지 않는 우승 쌍만 유지하여, 속성 메트릭을 향상시키면서 패턴을 보존합니다. 항목 및 패턴 홀드아웃 프로토콜 하에서 네 가지 속성 도메인을 사용하여 실험한 결과, 메트릭 기반 순서 학습된 DiT 모델이 사전 학습된 생성적 검색 모델보다 교차 지표 성능이 우수했으며, HPPO는 이를 더욱 향상시켰습니다. 8개의 도메인 분할 셀 중 7개에서 유의미한 개선 효과가 나타났으며, 가장 어려운 분할에서는 약간의 차이를 보였습니다. 메트릭 예측 검증, 순서 제거 실험, CPT/SFT 비교 및 후보 정책 제거 실험을 통해 성능 향상의 원인을 분석했습니다.

Original Abstract

Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scoring items. In many production settings this is not what is wanted: given a seed set that expresses a fine-grained pattern, one needs more items that both satisfy a target attribute and stay within that pattern. We formalize this as pattern-preserving attribute retrieval. The two goals pull against each other: averaging the seeds preserves the pattern but stays in a low-attribute region, while global attribute retrieval drifts to unrelated patterns. We approach the task with continuous generative retrieval, where a model reads a sequence of item embeddings and generates query embeddings for nearest-neighbor search. We propose MO-DiT+HPPO, a staged framework with raw-sequence pretraining, multi-domain metric-ordered continuation pretraining, tail-centroid fine-tuning, and HPPO. Metric-ordered training turns sparse online retrieval labels into in-pattern trajectories ordered from low to high predicted attribute density, teaching one model the metric-improvement direction across domains. HPPO aligns the generated query distribution with the true online objective by labeling a hybrid candidate pool with the online intersection metric and applying reference-anchored preference optimization. A Pareto pair filter keeps only winner pairs that do not lower same-pattern purity, raising the attribute metric without sacrificing the pattern. Across four attribute domains under item- and pattern-holdout protocols, metric-ordered DiT improves the intersection metric over a pretrained generative retriever, and HPPO improves it further, with significant gains on seven of eight domain-split cells and a marginal tie on the hardest split. Metric-predictor validation, order ablations, CPT/SFT comparisons, and a candidate-policy ablation show where the gains come from.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!