2608.04625v1 Aug 05, 2026 cs.AI

A/B 에이전트: 산업용 A/B 테스트를 위한 전략 반복을 위한 자기 진화 에이전트

A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing

Zhuohang Jiang
Zhuohang Jiang
Citations: 124
h-index: 3
Wenqi Fan
Wenqi Fan
Citations: 71
h-index: 3
Jun Wang
Jun Wang
Citations: 10
h-index: 1
Qing Li
Qing Li
Citations: 99
h-index: 5
Yuxin Chen
Yuxin Chen
Citations: 167
h-index: 6
Yongsen Pan
Yongsen Pan
Citations: 17
h-index: 3
Wenwu Ou
Wenwu Ou
Citations: 10
h-index: 1
Zheng Hu
Zheng Hu
Citations: 107
h-index: 6
Hongyang Wang
Hongyang Wang
Citations: 36
h-index: 3

산업용 추천 전략 반복은 대규모 A/B 실험에 크게 의존합니다. 기존 방식에서는 전문가가 전략을 반복적으로 설계하고, 실험을 구성하며, 결과를 분석하고, 매개변수를 조정해야 하므로 매우 노동 집약적이고 시간이 오래 걸립니다. 또한, 과거 실험에서 얻은 귀중한 지식은 종종 단편화되어 있어, 수동적인 전문가의 노력만으로는 체계적인 재사용이 어렵습니다. 기존 RAG 에이전트는 이전 전략을 검색하여 이러한 부담을 부분적으로 완화하지만, 일반적으로 경험을 평면적으로 구성하여 비즈니스 시나리오, 추천 단계, 최적화 목표 및 실험 환경 간의 계층적 관계를 고려하지 않습니다. 이는 종종 불일치하는 검색 결과를 초래하고 시나리오 간의 제한적인 전송을 야기하며, 에이전트가 순차적인 A/B 피드백을 통해 전략과 매개변수를 지속적으로 개선하는 것을 방해합니다. 이러한 한계를 해결하기 위해, 산업용 추천 전략 최적화를 위한 폐쇄 루프형 A/B 에이전트인 A/B 에이전트를 제안합니다. 이 프레임워크는 세 가지 핵심 구성 요소로 구성됩니다: 과거 전략 지식 조직화, 자율적인 목표 인식 전략 생성 및 실험 기반 전략 자기 진화. 이 시스템은 과거 전략을 계층적 경험 트리로 구성하고, 다중 경로 Tree-RAG를 통해 전송 가능한 증거를 검색하여 실행 가능한 전략을 생성하며, 온라인 A/B 피드백을 지속적으로 분석하여 자율적인 튜닝을 안내하고 경험 트리를 업데이트하여 자기 진화를 수행합니다. 광범위한 오프라인 및 온라인 평가 결과, 제안된 시스템은 실제 단편 동영상 전자 상거래 추천 시스템에서 GMV를 4.829% 향상시켰으며, 모든 안전 장치 지표에서 긍정적인 효과를 유지했습니다.

Original Abstract

Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repeatedly design strategies, configure experiments, analyze results, and adjust parameters, making the process labor-intensive and time-consuming. Meanwhile, valuable knowledge from historical experiments is often fragmented, making systematic reuse difficult through manual expert effort alone. Existing RAG agents partially alleviate this burden by retrieving prior strategies, but typically organize experience in a flat manner, overlooking the hierarchical relationships among business scenarios, recommendation stages, optimization objectives, and experimental contexts. This often results in mismatched retrieval and limited cross-scenario transfer, while preventing agents from continuously refining strategies and parameters through sequential A/B feedback. % To address these limitations, we propose A/B Agent, a closed-loop A/B agent for industrial recommendation strategy optimization. The framework comprises three tightly coupled core components: Historical Strategy Knowledge Organization, Autonomous Target-Aware Strategy Generation, and Experiment-Guided Strategy Self-Evolution. It organizes historical strategies into a hierarchical experience tree, retrieves transferable evidence through multi-path Tree-RAG to generate executable strategies, and continuously analyzes online A/B feedback to guide autonomous tuning and update the experience tree for self-evolution. Extensive offline and online evaluations demonstrate its effectiveness, including a 4.829% improvement in GMV in a real-world short-video e-commerce recommendation system while maintaining positive gains across all guardrail metrics.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!