2604.23993v1 Apr 27, 2026 cs.CL

EPM-RL: 전자상거래 환경에서의 온프레미스 제품 매핑을 위한 강화 학습

EPM-RL: Reinforcement Learning for On-Premise Product Mapping in E-Commerce

Minhyeong Yu
Minhyeong Yu
Citations: 5
h-index: 1
Wonduk Seo
Wonduk Seo
Citations: 26
h-index: 3

제품 매핑은 두 개의 전자상거래 상품 목록이 동일한 상품을 참조하는지 판단하는 핵심적인 문제로, 가격 모니터링 및 채널 가시성을 확보하는 데 중요한 역할을 합니다. 그러나 실제 마켓플레이스에서는 판매자들이 종종 프로모션 키워드, 플랫폼별 태그 및 번들 설명을 상품명에 추가하여 동일한 상품이 다양한 이름으로 나타나는 경우가 많습니다. 최근 LLM 기반 및 다중 에이전트 프레임워크는 이러한 어려운 경우에 대한 견고성과 해석 가능성을 향상시키지만, 종종 비용이 많이 드는 외부 API, 반복적인 검색 및 복잡한 추론 시간 오케스트레이션을 필요로 하므로, 대규모 배포가 어렵고 개인 정보 보호가 중요한 기업 환경에서 비용 효율적이지 않습니다. 이러한 문제를 해결하기 위해, 우리는 정확하고 효율적인 온프레미스 전자상거래 제품 매핑 모델을 구축하기 위한 강화 학습 기반 프레임워크인 EPM-RL을 제안합니다. 우리의 핵심 아이디어는 고비용 에이전트 기반 추론을 학습 가능한 내부 모델로 변환하는 것입니다. LLM이 생성한 설명과 인간의 검증을 거친 제품 쌍 세트에서 시작하여, 우리는 먼저 구조화된 추론 결과를 사용하여 작은 모델을 효율적으로 파인튜닝(PEFT)합니다. 그런 다음, 특별히 설계된 평가 모델로부터 얻은 출력 형식 준수, 레이블 정확성 및 추론-선호도 점수를 종합적으로 평가하는 에이전트 기반 보상을 사용하여 강화 학습(RL)을 통해 모델을 추가로 최적화합니다. 초기 결과는 EPM-RL이 PEFT만 사용한 학습보다 일관되게 성능이 향상되며, 상용 API 기반의 기존 모델보다 더 나은 품질-비용 균형을 제공하며, 동시에 개인 정보 보호를 보장하고 운영 비용을 절감할 수 있음을 보여줍니다. 이러한 결과는 강화 학습이 제품 매핑을 고지연 에이전트 기반 파이프라인에서 확장 가능하고 검사 가능하며 프로덕션 환경에 적합한 내부 시스템으로 전환할 수 있음을 시사합니다.

Original Abstract

Product mapping, the task of deciding whether two e-commerce listings refer to the same product, is a core problem for price monitoring and channel visibility. In real marketplaces, however, sellers frequently inject promotional keywords, platform-specific tags, and bundle descriptions into titles, causing the same product to appear under many different names. Recent LLM-based and multi-agent frameworks improve robustness and interpretability on such hard cases, but they often rely on expensive external APIs, repeated retrieval, and complex inference-time orchestration, making large-scale deployment costly and difficult in privacy-sensitive enterprise settings. To address these issues, we present EPM-RL, a reinforcement-learning-based framework for building an accurate and efficient on-premise e-commerce product mapping model. Our central idea is to distill high-cost agentic reasoning into a trainable in-house model. Starting from a curated set of product pairs with LLM-generated rationales and human verification, we first perform parameter-efficient fine-tuning (PEFT) on a small student model using structured reasoning outputs. We then further optimize the model with Reinforcement Learning (RL) using an agent-based reward that jointly evaluates output-format compliance, label correctness, reasoning--preference scores from specially designed judge models. Preliminary results show that EPM-RL consistently improves over PEFT-only training and offers a stronger quality--cost trade-off than commercial API-based baselines, while enabling private deployment and lower operational cost. These findings suggest that reinforcement learning can turn product mapping from a high-latency agentic pipeline into a scalable, inspectable, and production-ready in-house system.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!