2603.10400v1 Mar 11, 2026 cs.LG

텍스트 증거를 기반으로 한 서비스 시스템 설계

Designing Service Systems from Textual Evidence

Ruicheng Ao
Ruicheng Ao
Citations: 76
h-index: 5
David Simchi-Levi
David Simchi-Levi
Citations: 121
h-index: 6
Hongyu Chen
Hongyu Chen
Citations: 18
h-index: 3
Siyang Gao
Siyang Gao
Citations: 11
h-index: 1
Hanwei Li
Hanwei Li
Citations: 61
h-index: 4

서비스 시스템 설계는 대체 구성 요소 중에서 최적의 것을 선택하는 과정을 포함합니다. 여기에는 최적의 챗봇 변형, 최적의 라우팅 정책 또는 가장 효과적인 품질 관리 절차를 선택하는 것이 포함됩니다. 많은 서비스 시스템에서 성능 품질에 대한 주요 증거는 텍스트 형식으로 제공됩니다. 예를 들어 고객 지원 기록, 불만 사항 내용, 규정 준수 검토 보고서 등이 있습니다. 이러한 텍스트 증거를 대규모 언어 모델(LLM)이 분석하여 표준화된 품질 점수를 생성할 수 있지만, 이러한 자동화된 평가 도구는 대안 및 평가 사례에 따라 체계적인 편향을 나타냅니다. 인간 전문가의 검토는 정확하지만 비용이 많이 듭니다. 본 연구에서는 자동화된 평가가 저렴하지만 편향되어 있을 때, 비용이 많이 드는 인간 검토를 최소화하면서 가장 높은 신뢰도로 최적의 서비스 구성을 식별하는 방법을 연구합니다. 이를 위해 편향된 프록시 점수가 매 평가마다 관찰되고, 추가 비용을 들여 검증된 결과를 선택적으로 얻을 수 있는 순차적 의사 결정 문제로 formalize합니다. LLM만 사용할 경우 'arm-dependent bias' 하에서 최적의 선택이 불가능하며, 'naive selective-audit estimators'는 점근적으로 편향될 수 있음을 증명합니다. 우리는 프록시 점수에 역-경향성 가중 잔차를 결합한 추정기를 개발하고, 언제든지 유효한 신뢰 구간을 구성합니다. 우리의 알고리즘인 PP-LUCB는 평가할 대안을 동시에 결정하고, 인간 검토를 요청할지 여부를 결정하며, LLM 평가 도구가 신뢰도가 가장 낮은 곳에 검토를 집중합니다. 우리는 알고리즘의 정확성을 증명하고, 사례에 따라 달라지는 비용 경계를 설정하여 거의 최적의 효율성을 달성함을 보입니다. 고객 지원 티켓 분류 작업에서, 우리의 알고리즘은 40번의 실험에서 모두 최적의 모델을 정확하게 식별했으며, 동시에 감사 비용을 90% 절감했습니다.

Original Abstract

Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure. In many service systems, the primary evidence of performance quality is textual -- customer support transcripts, complaint narratives, compliance review reports -- rather than the scalar measurements assumed by classical optimization methods. Large language models (LLMs) can read such textual evidence and produce standardized quality scores, but these automated judges exhibit systematic biases that vary across alternatives and evaluation instances. Human expert review remains accurate but costly. We study how to identify the best service configuration with high confidence while minimizing expensive human audits, given that automated evaluation is cheap but biased. We formalize this as a sequential decision problem where a biased proxy score is observed for every evaluation, and a verified outcome can be acquired selectively at additional cost. We prove that LLM-only selection fails under arm-dependent bias, and that naive selective-audit estimators can be asymptotically biased. We develop an estimator combining proxy scores with inverse-propensity-weighted residuals and construct anytime-valid confidence sequences. Our algorithm, PP-LUCB, jointly decides which alternatives to evaluate and whether to request human audits, concentrating reviews where the LLM judge is least reliable. We prove correctness and establish instance-dependent cost bounds showing near-optimal efficiency. On a customer support ticket classification task, our algorithm correctly identifies the best model in 40/40 trials while achieving 90\% audit cost reduction.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!