AI 평가를 위한 '사과와 사과' 비교를 향하여: 실제 사용 사례에서 평가 시나리오까지
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
AI 시스템 비교를 위한 다양한 방법론과 측정 지표는 종종 AI 평가에서 '사과와 오렌지'를 비교하는 듯한 인상을 줍니다. 본 연구는 실제 AI 평가에서 '사과와 사과' 비교를 가능하게 하기 위해 평가 시나리오의 방법론적 투명성, 실제 적용 가능성, 그리고 인간 중심 설계(HCD) 원칙의 중요성을 강조합니다. 본 연구에서는 주제 전문가(SMEs)로부터 사용 사례를 수집하고, 6가지 핵심 요소(사용 사례, 산업 분야, 사용자(직접 및 간접), 예상되는 결과, 예상되는 영향(긍정 및 부정), 핵심 성과 지표(KPI) 및 측정 지표)를 포함한 체계적인 AI 사용 사례 워크시트를 사용하여, 고수준 사용 사례를 상세한 시나리오로 변환하는 반복 가능한 프로세스를 제안합니다. 본 연구는 미국 금융 서비스 부문에서 워크시트 및 프로세스의 유용성을 입증합니다. 본 논문에서는 금융 서비스 부문 SMEs가 식별한 고수준 AI 사용 사례의 예시를 제시합니다. 여기에는 사이버 방어 강화, 개발자 생산성 향상, 금융 범죄 통합, 의심스러운 활동 보고서(SAR) 작성, 신용 메모 생성, 그리고 내부 콜센터 지원 등이 포함됩니다. 제시된 AI 사용 사례는 예시이며, 전체를 포괄하지 않습니다. 본 연구의 핵심은 LLM 프롬프팅과 인간 검토를 결합한 3단계 확장 파이프라인을 사용하여 SMEs로부터 수집된 사용 사례로부터 107개의 시나리오를 생성하는 것입니다. 이 프로세스는 시나리오 제목 및 설명, 사용자, 이점 및 위험, 측정 지표, 그리고 시나리오 내용 및 평가 목표와 같은 핵심 시나리오 요소에 대한 반복적인 인간 검토를 통해 실제 적용 가능성을 보장합니다. 인간 검토는 시나리오가 실제 사용 및 인간의 요구 사항을 반영하는지 확인합니다. 본 연구에서는 시나리오 품질을 평가하기 위한 검증 기준을 제시합니다. 본 연구는 핵심 시나리오 구성 요소를 정의함으로써 인간 중심 AI 평가를 위한 보다 일관되고 의미 있는 패러다임을 지원합니다.
AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be "apples-to-oranges" comparisons across AI evaluations. To move toward "apples-to-apples" comparisons in real-world AI evaluations, this work advocates for methodological transparency in evaluation scenarios, operational grounding, and human-centered design (HCD) principles. We propose a repeatable process for transforming high-level use cases to detailed scenarios by eliciting use cases from subject matter experts (SMEs) via a structured AI Use Case Worksheet with six key elements: use case, sector, user (direct and indirect), intended outcomes, expected impacts (positive and negative), and KPIs and metrics. We demonstrate utility of the worksheet and process in the U.S. financial services sector. This paper reports on example high-level AI use cases identified by financial services sector SMEs: cyber defense enablement, developer productivity, financial crime aggregation, suspicious activity report (SAR) filing, credit memo generation, and internal call center support. These AI use cases provided are illustrative of the process and not exhaustive. Central to our work is a three-stage expansion pipeline combining LLM prompting with human reviews to generate 107 scenarios from those use cases elicited from SMEs. This process integrates iterative human reviews at every juncture to ensure operational grounding: for scenario titles and descriptions; for core scenario elements like users, benefits and risks, and metrics; and for scenario narratives and evaluation objectives. Human checkpoints ensure scenarios remain reflective of real-world usage and human needs. We describe a validation rubric to assess scenario quality. By defining key scenario components, this work supports a more consistent and meaningful paradigm for human-centered AI evaluations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.