Litmus: AI 시스템 평가를 위한 코드 기반의 제로-레이블 메트릭 명세
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
에이전트 기능을 가진 LLM 시스템이 프로토타입 단계를 넘어 다양한 분야로 확장되면서, 이러한 시스템을 평가하는 것이 더욱 중요해지고 어렵게 되었습니다. 문제는 개별적인 지표의 신뢰성이 떨어질 뿐만 아니라, 평가 목표가 종종 명시적으로 정의되지 않는다는 점입니다. 어떤 시스템이 무엇을 수행해야 하는지, 어떻게 실패할 수 있는지, 그리고 어떤 종류의 실패가 중요한지에 대한 명확한 설명 없이, 지표 선택은 정당화하거나 해석하고 검증하기 어렵습니다. 본 논문에서는 Litmus라는 제로-레이블 시스템을 소개합니다. Litmus는 소스 코드에서 평가 의도를 추출하고, 목표 지향적인 질문을 통해 AI 파이프라인의 평가 및 모니터링 지표를 설계합니다. Litmus는 평가 대상이 이미 알려져 있다고 가정하는 대신, 무엇을 측정해야 하고 왜 그래야 하는지를 먼저 파악한 다음, 이러한 답변을 정당화된 단계별 지표 포트폴리오를 구성하기 위한 제약 조건으로 변환합니다. 우리는 Litmus를 금융 계정 그룹화, 과학 QA 및 잠재적 위험 평가라는 세 가지 실제 코드 기반 AI 파이프라인에서 AutoMetrics 및 세 가지 DynamicRubric 기준과 비교하여 평가했습니다. Litmus는 가장 넓거나 동등하게 넓은 범위의 고려 사항을 포괄하고, 더 많은 파이프라인 단계를 포함하며, 거의 중복되지 않는 지표 포트폴리오를 생성하고, 세 가지 파이프라인 모두에서 개별 행에 대한 품질 레이블에 대한 유효성 측면에서 가장 높은 순위를 차지했습니다. 특히 과학 QA 분야에서는 Spearman 상관 계수가 0.72로 나타났으며, 이는 모든 기준의 경우 0.47 미만인 반면, 라벨을 사용하지 않고 지표를 설계했음에도 불구하고 감사 프레임워크의 두 가지 구성 요소와 비교하여 중첩되는 신뢰 구간 내에 있었습니다. 우리의 결과는 자동 지표 구현에서 자동 지표 명세로의 전환을 지원합니다. 어떤 지표를 계산해야 할지 묻기 전에, 평가 시스템은 무엇을 측정해야 하는지와 왜 그래야 하는지를 먼저 질문해야 합니다.
As agentic LLM systems move from prototypes to deployment across increasingly diverse domains, evaluating them has become both more important and more difficult. The challenge is not only that individual metrics may be unreliable, but that evaluation goals are often left implicit. Without a clear account of what a system is expected to do, how it can fail, and which failures matter, metric choices become difficult to justify, interpret, or validate. We present Litmus, a zero-label system that designs evaluation and monitoring metrics for AI pipelines by eliciting evaluation intent from source code and targeted interrogation. Instead of assuming that the evaluation target is already known, Litmus first identifies what must be measured and why, then converts those answers into constraints for constructing a justified, per-stage metric portfolio. We evaluate Litmus on three real, code-defined AI pipelines - financial account grouping, scientific QA, and inherent risk assessment - against AutoMetrics and three DynamicRubric baselines. Litmus achieves the broadest or tied-broadest concern coverage, spans more pipeline stages, produces a near-zero-redundancy portfolio, and ranks first in validity against per-row quality labels on all three pipelines - decisively on scientific QA (Spearman $ρ=0.72$ vs. less than $0.47$ for every baseline), and within overlapping confidence intervals in relation to two components of the audit framework despite using no labels during metric design. Our results support a shift from automatic metric implementation to automatic metric specification: before asking which metric to compute, evaluation systems should ask what must be measured and why.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.