2607.29213v1 Jul 31, 2026 cs.IR

GALA: 태오바오 상구 추천 시스템에서의 적응형 다중 모드 표현을 위한 생성적 정렬 학습

GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

Zisen Sang
Zisen Sang
Citations: 12
h-index: 2
Guodong Cao
Guodong Cao
Citations: 15
h-index: 2
Jia Jia
Jia Jia
Citations: 39
h-index: 4
Jiping Liu
Jiping Liu
Citations: 0
h-index: 0
Zhongmin Zhang
Zhongmin Zhang
Citations: 0
h-index: 0
Zhijia Fang
Zhijia Fang
Citations: 0
h-index: 0
Ouyang Tao
Ouyang Tao
Citations: 1,001
h-index: 12
Ma Jiang
Ma Jiang
Citations: 0
h-index: 0
Shaopeng Liang
Shaopeng Liang
Citations: 0
h-index: 0
Zeyang Hou
Zeyang Hou
Citations: 0
h-index: 0

최신 음식 배달 추천 시스템은 사용자 경험 향상을 위해 이미지, 텍스트 및 사용자 인터랙션 기록과 같은 다양한 정보를 활용하지만, 이러한 이질적인 정보들을 효과적으로 통합하는 것은 여전히 어려운 과제이며, 이는 다중 모드 정보의 통합 모델링과 변화하는 사용자 의도에 대한 적응을 저해합니다. 기존의 주요 두 단계 접근 방식에서 이미지-텍스트 인코더의 내용-의미 사전 학습과 사용자 행동 기반 랭킹 모델 간의 분리는 의미 이해와 사용자 행동 패턴 간의 정렬을 제한합니다. 이러한 문제를 해결하기 위해, 우리는 핵심적인 혁신이 사용자의 행동으로부터 다중 모드 사전 학습 데이터를 생성하고, 전환 기반 보상을 통해 이를 개선하는

Original Abstract

Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline whose core innovation lies in an intermediate "generative RL alignment" stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. GALA comprises three stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, a novel intermediate stage that refines multimodal embeddings through reward-driven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55 percent increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!