다중 채널 업리프트 정책 학습
Multi-channel Uplift Policy Learning
전자상거래 플랫폼은 제한된 마케팅 예산을 여러 채널에 할당하여 비즈니스 효용을 극대화해야 합니다. 그러나 기존의 예측-최적화(PTO) 패러다임은 관찰 데이터의 교란 효과 및 심각한 외삽 문제로 인해 이러한 복잡한 환경에서 제대로 작동하지 않습니다. 본 연구에서는 이 문제를 시플렉스 제약 조건 하에서의 업리프트 의사 결정 문제로 정의하고, 빠른-느린 인과 추론 프레임워크인 ReAlloc을 제안합니다. 구체적으로, 민첩한 Orthogonal Teacher는 단기 로그 데이터로부터 편향되지 않은 지역적 기울기를 추출하며, Explanation-Guided Student는 이를 장기적인 관점에서 구조화된 주변 분포로 변환합니다. 이러한 설계는 채널 간의 상호 영향을 고려하는 보수적이고 신중한 의사 결정을 가능하게 합니다. 광범위한 시뮬레이션과 Taobao 플랫폼에서의 대규모 온라인 A/B 테스트를 통해 ReAlloc이 지불 주문 수와 수익 모두에서 동시에 향상을 달성한다는 것을 확인했습니다.
E-commerce platforms must allocate fixed marketing budgets across multiple channels to maximize business utility. However, standard predict-then-optimize (PTO) paradigms fail in this compositional space due to observational confounding and severe extrapolation. We formulate this challenge as a simplex-constrained uplift decision problem and propose ReAlloc, a fast-slow causal framework. Specifically, an agile Orthogonal Teacher extracts unbiased local gradients from short-term logs, while an Explanation-Guided Student distills them into a structured marginal field over long-term horizons. This design enables support-aware, conservative decisions that capture cross-channel substitutions. Extensive simulations and large-scale online A/B tests on Taobao platform demonstrate that ReAlloc achieves simultaneous lifts in both pay order and income.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.