SynthAVE: LLM 아레나 검증을 통한 대규모 전자상거래 합성 레이블링
SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
대규모 언어 모델(LLM)을 사용하여 전자상거래 제품 속성 추출을 위해 수천 가지의 제품 유형, 속성 및 여러 언어에 걸쳐 대표적인 레이블이 지정된 데이터가 필요합니다. 이러한 조합은 수백만 개의 어노테이션으로 이어지며, 이는 인적 레이블링으로는 감당하기 어렵습니다. 최근 연구에서는 LLM을 사용하여 합성 레이블을 생성하는 방법을 보여주었지만, 이러한 접근 방식을 산업 규모로 적용하려면 통합된 품질 관리 메커니즘이 필요합니다. 본 논문에서는 12,726개의 제품(229개 제품 유형, 792개 속성, 4개 언어: 스페인어, 프랑스어, 이탈리아어, 독일어)에 대한 속성 값 추출을 위한 대규모의 인간 검증 벤치마크인 SynthAVE를 제시합니다. 합성 레이블을 대규모로 검증하기 위해, 21개의 평가 구성(7개 모델 패밀리 $ imes$ 3가지 프롬프트)이 각 샘플을 독립적으로 평가하는 다중 LLM 아레나 프레임워크를 도입했습니다. 최종 레이블은 다수결 투표를 통해 결정됩니다. 다수결 투표 집합은 인간 전문가와 Cohen's $κ= 0.92$ (95.2% 일치)의 높은 합의도를 보이며, 개별 평가자는 Fleiss' $κ= 0.76$의 상당한 모델 간 일관성을 나타냅니다. 이는 다양한 모델이 개별적인 판단을 내렸지만, 이러한 판단들이 결합되어 매우 신뢰할 수 있는 예측 결과를 제공하며, 인간 검토와 동등한 품질을 유지하면서 비용 효율적인 대규모 검증을 가능하게 함을 보여줍니다.
Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs, deploying such approaches at industrial scale requires integrated quality control mechanisms. We present SynthAVE, a large-scale human-validated benchmark for attribute value extraction spanning 12,726 products across 229 product types, 792 attributes, and 4 languages (Spanish, French, Italian, German). To validate synthetic labels at scale, we introduce a multi-LLM arena framework where samples are independently evaluated by 21 judge configurations (7 model families $\times$ 3 prompts), with final labels determined via majority voting. The majority vote ensemble agrees with human experts at Cohen's $κ= 0.92$ (95.2% agreement), while individual judges show substantial inter-model agreement (Fleiss' $κ= 0.76$). This demonstrates that diverse models with varying individual judgments aggregate into highly reliable predictions, enabling cost-effective validation at scale while maintaining quality parity with human review.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.