전역 수준 제약 조건 하에서의 인센티브 광고를 위한 생성적 최적화
Generative Optimization for Incentivized Advertising with Global Level Constraints
인센티브 광고는 사용자 참여를 유도하기 위해 금전적 또는 가상 보상을 제공하며, 주요 과제는 엄격한 전역 제약 조건 하에서 연속적인 인센티브 크기를 최적화하는 것입니다. 이 문제는 높은 빈도의 상호 작용, 지연된 피드백 및 사용자의 피로와 같은 비마르코프 특성으로 인해 기존의 효과 측정 모델링 및 제약 강화 학습 접근 방식의 효과를 제한합니다. 이러한 과제를 해결하기 위해, 우리는 시스템 수준의 전역적인 제약을 고려하는 생성적 프레임워크인 GOAL을 제안합니다. GOAL은 인센티브 할당을 조건부 시퀀스 생성 문제로 공식화하며, 사용자 이력 및 시스템 레벨의 전반적인 압박에 따라 직접적으로 인센티브 크기를 생성하고, 계층적 원인 관계 상태 인코더를 통합하여 지역적인 행동 동역학과 장기 의존성을 모두 포착합니다. 유연한 제약 조건 제어를 가능하게 하기 위해, 우리는 ROI 제약 조건의 스펙트럼에 걸쳐 일반화되는 단일 생성 정책을 학습하는 extbf{S}afe extbf{C}onstrained extbf{P}olicy extbf{O}ptimization (SCPO)를 도입합니다. 대규모 실제 데이터 및 피로를 고려한 합성 환경에서의 실험 결과, GOAL은 장기적인 수익과 사용자 유지율을 향상시키고 강력한 기준 모델에 비해 ROI 위반률을 크게 줄이는 것으로 나타났습니다.
Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes under strict global constraints. This problem is complicated by high-frequency interactions, delayed feedback, and non-Markovian user dynamics such as fatigue, which limit the effectiveness of existing uplift modeling and constrained reinforcement learning approaches. To address these challenges, we propose GOAL, a constraint-aware generative framework that formulates incentive allocation as a conditional sequence generation problem. GOAL directly generates incentive magnitudes conditioned on user histories and system-level global pressure, and integrates a hierarchical causal state encoder to capture both local behavioral dynamics and long-range dependencies. To enable flexible constraint control, we introduce \textbf{S}afe \textbf{C}onstrained \textbf{P}olicy \textbf{O}ptimization (SCPO), which learns a single generative policy that generalizes across a spectrum of ROI constraints without retraining. Experiments on large-scale real-world data and a synthetic fatigue-aware environment show that GOAL improves long-term revenue and user retention while substantially reducing ROI violation rates compared to strong baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.