2603.24226v4 Mar 25, 2026 cs.IR

UniScale: 검색 순위를 위한 데이터 및 모델의 시너지 효과를 통한 전체 공간 확장

UniScale: Synergistic Entire Space Data and Model Scaling for Search Ranking

Dan Ou
Dan Ou
Citations: 22
h-index: 2
Haihong Tang
Haihong Tang
Citations: 52
h-index: 4
Liren Yu
Liren Yu
Citations: 10
h-index: 1
Tao Zhang
Tao Zhang
Citations: 36
h-index: 1
Caiyuan Li
Caiyuan Li
Citations: 106
h-index: 1
Feiyi Dong
Feiyi Dong
Citations: 4
h-index: 1
Zhixuan Zhang
Zhixuan Zhang
Citations: 1
h-index: 1
Bo Zheng
Bo Zheng
Citations: 13
h-index: 2

최근 대규모 언어 모델(LLM)의 발전은 산업용 검색, 광고 및 추천 시스템 분야에서 스케일링 연구에 대한 관심을 크게 높였습니다. 그러나 기존 접근 방식은 주로 아키텍처 개선에 초점을 맞추고 있으며, 데이터 설계와 아키텍처 간의 중요한 시너지 효과를 간과하고 있습니다. 우리는 모델 파라미터만 확장하는 것이 투자 대비 효과가 감소하며, 복잡한 이질적인 데이터 분포로 인한 성능 저하가 모델 설계만으로는 회복하기 어려운 경우가 많다는 것을 확인했습니다. 본 논문에서는 데이터와 아키텍처를 공동으로 최적화하여 모델 스케일링의 잠재력을 최대한 발휘하는 새로운 프레임워크인 UniScale을 제안합니다. UniScale은 다음 두 가지 핵심 부분으로 구성됩니다: (1) ES$^3$ (전체 공간 샘플링 시스템): 계층적 레이블 할당과 크로스 도메인 검색화를 통해 기존 샘플링 전략을 넘어 학습 신호를 확장하는 고품질 데이터 스케일링 시스템입니다. (2) HHSFT (이질적인 계층적 샘플 융합 트랜스포머): 이질적인 계층적 특징 상호 작용과 전체 공간 사용자 관심사 융합을 통해 확장된 데이터의 복잡하고 이질적인 분포를 효과적으로 모델링하는 새로운 아키텍처로, 구조만으로 모델을 조정하는 것의 성능 한계를 뛰어넘습니다. 대규모 산업 데이터 세트에 대한 광범위한 실험 결과, UniScale은 상당한 성능 향상을 달성하며 명확한 스케일링 추세를 보여줍니다. 실제 전자 상거래 검색 플랫폼에서의 온라인 A/B 테스트는 UniScale이 강력한 기존 시스템을 꾸준히 능가하며, 사용자 구매율과 총 상품 가치(GMV)를 각각 1.70% 및 2.04% 증가시키는 것을 확인했습니다.

Original Abstract

Recent advances in Large Language Models (LLMs) have inspired a surge of scaling research in industrial search, advertising, and recommendation systems. However, existing approaches focus mainly on architectural improvements, overlooking the critical synergy between data and architecture design. We observe that scaling model parameters alone exhibits diminishing returns, and that the performance degradation caused by complex heterogeneous data distributions is often irrecoverable through model design alone. In this paper, we propose UniScale, a novel co-design framework that jointly optimizes data and architecture to unlock the full potential of model scaling. UniScale includes two core parts: (1) ES$^3$ (Entire-Space Sample System), a high-quality data scaling system that expands the training signal beyond conventional sampling strategies through intra-domain expansion with hierarchical label attribution and cross-domain searchification; and (2) HHSFT (Heterogeneous Hierarchical Sample Fusion Transformer), a novel architecture that effectively models the complex heterogeneous distribution of scaled data via Heterogeneous Hierarchical Feature Interaction and Entire Space User Interest Fusion, thereby surpassing the performance ceiling of structure-only model tuning. Extensive experiments on large-scale industrial datasets demonstrate that UniScale achieves significant improvements and exhibits clear scaling trends. Online A/B tests on a real-world E-commerce search platform confirm that UniScale consistently outperforms strong production baselines, achieving 1.70% and 2.04% increases in user purchase and Gross Merchandise Value (GMV).

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!