2606.18803v1 Jun 17, 2026 cs.AI

ProfiLLM: 산업용 택시 호출 시스템의 효율성 향상을 위한 사용자 프로파일링

ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch

Hao Liu
Hao Liu
Citations: 125
h-index: 5
Tengfei Lyu
Tengfei Lyu
Citations: 10
h-index: 2
Zirui Yuan
Zirui Yuan
Citations: 31
h-index: 4
Xu Liu
Xu Liu
Citations: 0
h-index: 0
Kaiyang Wan
Kaiyang Wan
Citations: 23
h-index: 4
Zihao Lu
Zihao Lu
Citations: 19
h-index: 2
Li Ma
Li Ma
Citations: 2
h-index: 1

대규모 언어 모델(LLM)을 활용하여 플랫폼 전체의 행동 로그에서 의미론적 특징을 추출하는 것은 산업용 택시 호출 시스템 운영에 있어 매력적인 동시에 아직 덜 연구된 데이터 시스템 문제입니다. 현재 생산 환경에서의 매칭 파이프라인은 주로 구조화된 숫자형 특징에 의존하고 있지만, 중요한 행동 패턴(예: 특정 지역에 대한 운전자의 습관적인 회피)은 본질적으로 문맥적이며 LLM을 통해 생성된 사용자 프로필로 자연스럽게 표현될 수 있습니다. 그러나 이러한 프로파일링을 실시간으로, 밀리초 수준의 지연 시간 내에서 운영하는 것은 다음과 같은 세 가지 복잡하게 얽힌 제약 조건에 직면합니다. 첫째, 플랫폼에서 매일 발생하는 수백만 건의 주문 데이터는 어떤 LLM의 컨텍스트 창보다 훨씬 방대합니다. 둘째, 대부분의 사용자는 상위 사용자보다 훨씬 적은 상호 작용을 가지므로 개별 사용자 프로파일링이 어렵습니다. 셋째, 표면적으로 유용한 프로필이 반드시 다운스트림 예측 성능 향상으로 이어지지 않을 수 있습니다. 본 연구에서는 ProfiLLM이라는 LLM 기반 데이터 파이프라인을 제안하며, 이는 생산 매칭 시스템에서 효율성을 고려한 사용자 프로파일링을 구현하기 위해 두 가지 모듈로 구성됩니다. (1) 도구 강화 글로벌 지식 마이닝: LLM 에이전트에게 27가지 분석 도구를 제공하여 플랫폼 전체 데이터를 분석하고 재사용 가능한 글로벌 지식을 생성하며, 사용자 클러스터링 규칙과 지역 수준의 공급-수요 예측 정보를 추출합니다. (2) 효율성 기반 프로필 탐색: 각 클러스터별로 여러 후보 프로필을 생성하고, 경량화된 다운스트림 유틸리티 프록시를 통해 평가하며, 가장 우수한 후보 프로필을 반복적으로 개선하고 DPO(Direct Preference Optimization) 파인튜닝을 위한 선호도 쌍을 구성합니다. DiDi의 실제 택시 호출 시스템에 ProfiLLM을 적용한 결과, 예측 정확도가 최대 +6.14% 향상되었고, 시뮬레이션 기반 배차 성능은 최대 +4.35% 개선되었습니다. 또한 14일 동안 진행된 온라인 A/B 테스트에서 GMV(Gross Merchandise Value)가 +0.47%, 완료율이 +0.33%, 수락 전 취소율이 -0.82% 향상되는 등 일관성 있는 성능 향상을 확인했습니다.

Original Abstract

Bringing Large Language Models (LLMs) into industrial ride-hailing dispatch as semantic feature extractors over platform-scale behavioral logs is a compelling but under-explored data systems problem. Production matching pipelines remain dominated by structured numerical features, yet decisive behavioral signals (e.g., a driver's habitual aversion to certain regions) are inherently contextual and naturally expressible as LLM-generated user profiles. However, scaling such profiling to a live, millisecond-latency dispatcher faces three intertwined constraints rarely addressed together: on a platform with millions of daily orders, logs exceed any LLM's context window by orders of magnitude; most users are long-tail, with too few interactions for per-user profiling; and surface-fluent profiles do not necessarily improve downstream prediction utility. We present ProfiLLM, an agentic LLM data pipeline that operationalizes utility-aligned user profiling for production matching systems through two modules. (1) Tool-Augmented Global Knowledge Mining equips an LLM agent with 27 analytical tools to mine platform-scale data, producing reusable global knowledge, adaptive user clustering rules, and region-level supply-demand priors. (2) Utility-Aligned Profile Exploration generates multiple candidate profiles per cluster, evaluates them via a lightweight downstream utility proxy, iteratively refines the best candidates and constructs preference pairs for DPO fine-tuning. Deployed on DiDi's production dispatcher, ProfiLLM achieves up to +6.14% relative AUC improvement in outcome prediction, up to +4.35% GMV gain in dispatching simulation, and consistent improvements in a 14-day online A/B test including +0.47% GMV, +0.33% Completion Rate, and -0.82% Cancel-Before-Accept rate.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!