프로젝트 칼레이도스코프: 실제 AI 애플리케이션을 위한 문맥 기반, 인간 중심 평가
Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
평가(Eval)는 실제 AI 애플리케이션의 배포 과정에서 중요한 병목 현상입니다. 공개 벤치마크는 종종 특정 팀의 사용자, 환경 또는 정책과 일치하지 않으며, 수동 검토는 확장하기 어렵습니다. 본 연구는 공공 부문 AI 애플리케이션 관련 작업 경험을 바탕으로, 애플리케이션이 지역 정책 및 규제 요구 사항을 충족해야 하는 경우 발생하는 반복적인 평가 문제를 해결하고자 합니다. 우리는 문맥 기반 기능 평가를 위한 통합 워크플로우인 칼레이도스코프를 제시합니다. 칼레이도스코프는 페르소나 기반 테스트 생성, 문맥화된 평가 기준, 그리고 신뢰성 확보를 위한 자동 점수 시스템을 결합한 인간 검토 프로세스를 포함합니다. 생성된 테스트 케이스는 애플리케이션별 평가 기준에 따라 점수가 매겨지며, 인간의 주석은 검토 가능한 레이블을 제공하고, LLM(대규모 언어 모델) 판사는 인간의 판단과 일치하는 수준이 설정된 임계값을 충족할 때만 자동 점수를 수행합니다. 따라서 칼레이도스코프는 제품 팀에게 실용적이고 투명하며 반복적인 워크플로우를 제공합니다. 본 연구에서는 4개의 조직 사용 사례에 대한 3주간의 시범 운영 결과와, 108개의 주석이 달린 질의응답 쌍을 활용한 4가지 영역 및 14가지 평가 차원에 대한 사용자 정의 평가 기준 실험 결과를 제시합니다. 이러한 결과는 전체적인 신뢰성 있는 자동 점수 시스템 구축에 유용한 기능을 보여줍니다.
Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.