2606.24605v1 Jun 23, 2026 cs.AI

ScaleToT: 대규모 저 활동 사용자 모델링을 위한 구조화된 LLM 추론의 일반화

ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling

Han Li
Han Li
Citations: 20
h-index: 3
Kun Gai
Kun Gai
Citations: 1,215
h-index: 12
Linxun Chen
Linxun Chen
Citations: 207
h-index: 9
Tianbao Ma
Tianbao Ma
Citations: 5
h-index: 2
Zhaojie Liu
Zhaojie Liu
Citations: 141
h-index: 7
Yanan Niu
Yanan Niu
Citations: 43
h-index: 4
Changyuan Xi
Changyuan Xi
Citations: 1
h-index: 1
Yichuan Zou
Yichuan Zou
Citations: 5
h-index: 2
Chengen Li
Chengen Li
Citations: 3
h-index: 1
Zilong Lu
Zilong Lu
Citations: 108
h-index: 4

정확한 사용자 모델링은 종종 풍부한 상호 작용 기록에 의존하지만, 이는 수십억 명의 저 활동 사용자에 대해서는 사용할 수 없습니다. 대규모 언어 모델(LLM)은 정적 프로필에서 잠재적인 사용자 상태를 추론할 수 있지만, 프로필이 희소하면 이러한 추론이 신뢰성이 떨어지며, LLM을 수십억 명의 사용자에게 적용하는 것은 비용 효율적이지 않습니다. 본 논문에서는 소규모 LLM으로 처리된 일부 데이터를 기반으로 구조화된 추론을 학습하고 이를 더 넓은 범위의 저 활동 사용자 집단에 확장하는 ScaleToT를 제안합니다. ScaleToT는 추론의 신뢰성을 향상시키기 위해, 제한된 엔트로피 기반 트리-오브-생트(Tree-of-Thought, ToT) 개선 절차를 통해 유형화된 사용자 상태 체인을 구성합니다. 또한, 희소한 프로필에서도 이러한 구조화된 추론을 활용할 수 있도록, 전문가가 큐레이션한 체인을 사용하여 지도 학습(Supervised Fine-Tuning, SFT) 및 결과 기반 세그먼트 인식 암묵적 보상 정책 최적화(Outcome-Driven Segment-Aware Implicit Reward Policy Optimization, OSIPO)를 통해 정적 프로필에 대한 학생 모델을 학습시킵니다. ScaleToT는 이후 학생 모델의 추론 표현을 경량 프로파일 인코더로 이전하여 LLM 추론 없이도 나머지 사용자에게 공유된 추론 신호를 제공합니다. 본 논문에서는 수십억 규모의 광고 환경에서 생애 가치(Lifetime Value, LTV) 예측에 대한 ScaleToT의 성능을 평가했습니다. 무작위 온라인 A/B 테스트 결과, LT30이 6.738% 증가했으며, 오프라인 추론은 잠재적 사용자 집단의 7.32%에 불과하여 전체 사용자 기반에 대한 추론에 비해 컴퓨팅 비용을 크게 절감할 수 있었습니다.

Original Abstract

Accurate user modeling often depends on rich interaction histories, which are unavailable for billions of low-activity users. Large Language Models (LLMs) can infer latent user states from static profiles, but this reasoning becomes unreliable when profiles are sparse, and applying an LLM to billions of users is prohibitively expensive. We present ScaleToT, which learns structured reasoning from a small LLM-processed subset and extends it to the broader low-activity user population. To improve reasoning reliability, ScaleToT constructs typed user-state chains with a bounded entropy-guided Tree-of-Thought (ToT) refinement procedure. To make this structured reasoning usable from sparse profiles, the teacher-curated chains are used to train a student model on static profiles through supervised fine-tuning (SFT) and Outcome-Driven Segment-Aware Implicit Reward Policy Optimization (OSIPO). ScaleToT then transfers the student's reasoning representations to a lightweight profile encoder, providing shared reasoning signals for the remaining users without LLM inference. We evaluate ScaleToT on lifetime value (LTV) prediction in a billion-scale advertising deployment. A randomized online A/B test increased LT30 by 6.738\%, while offline reasoning covered only 7.32\% of the potential population, greatly reducing compute cost compared with full-population reasoning.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!