2605.26001v1 May 25, 2026 cs.CL

AI 지원 체계화 방법론을 활용한 생성형 AI 시스템 평가

AI-Assisted Systematization for Evaluating GenAI Systems

Hussein Mozannar
Hussein Mozannar
Microsoft
Citations: 1,700
h-index: 18
Solon Barocas
Solon Barocas
Citations: 11,463
h-index: 39
Alexandra Chouldechova
Alexandra Chouldechova
Citations: 205
h-index: 7
Dhruv Agarwal
Dhruv Agarwal
Citations: 166
h-index: 5
Emily Sheng
Emily Sheng
Citations: 132
h-index: 6
C. Atalla
C. Atalla
Citations: 190
h-index: 7
J. Garcia-Gathright
J. Garcia-Gathright
Citations: 672
h-index: 15
Hannah Washington
Hannah Washington
Citations: 100
h-index: 4
Hanna M. Wallach
Hanna M. Wallach
Citations: 242
h-index: 7

생성형 AI (GenAI) 시스템 평가는 "추론", "공정성", 또는 "창의성"과 같이 광범위하고 논쟁적인 개념들을 평가 대상으로 삼기 때문에 어려움을 겪습니다. 이러한 개념들이 명확하게 정의되지 않으면, 무엇을 측정해야 하는지, 그리고 어떻게 평가 결과를 해석해야 하는지에 대한 불확실성이 발생합니다. 이 문제는 체계화 과정의 부재에서 비롯됩니다. 즉, 광범위한 배경 개념을 구체적이고 구조화된 방식으로 측정 가능한 형태로 변환하는 과정이 누락된 것입니다. 본 연구는 이러한 체계화 과정이 인지적으로 많은 부담과 자원을 필요로 한다는 점을 고려하여, AI 지원이 이 과정을 어떻게 보조할 수 있는지 조사합니다. AI 지원 체계화를 가능하게 하고 그 품질을 평가하기 위해, 우리는 체계화된 개념의 구조적 표현인 "개념 사양(concept spec)" 및 검증 워크시트를 제안합니다. 또한, 직접적인 제로샷 접근 방식과 기존 문헌의 수동 체계화 방식을 더욱 가깝게 반영하는 멀티 에이전트 접근 방식을 포함하여 두 가지 AI 지원 체계화 도구를 개발했습니다. 우리는 이러한 체계화 도구를 사용하여 "증오 기반 담론" 및 "디지털 공감"이라는 두 가지 개념에 대한 개념 사양을 생성하고, 결과적으로 생성된 개념 사양의 내용 타당성 및 정보 검색 가능성을 평가했습니다.

Original Abstract

Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as "reasoning," "fairness," or "creativity." When these concepts are left underspecified, it becomes unclear what should be measured or how evaluation results should be interpreted. This problem reflects a missing step: systematization, that is, moving from a broad background concept to an explicit, structured account of the concept in measurable terms. To help address the fact that systematization is cognitively demanding and resource-intensive, we investigate whether AI assistance can support this process. To enable AI-assisted systematization and assess its quality, we introduce a structured representation of a systematized concept, a concept spec, and a validation worksheet. We then develop two AI-assisted systematizers: a direct, zero-shot approach and a multi-agent approach that more closely mirrors manual systematization approaches from existing literature. We use these systematizers to produce concept specs for two concepts -- hate-based rhetoric and digital empathy -- and evaluate resulting concept specs on content validity and information recoverability.

0 Citations
0 Influential
19.5 Altmetric
97.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!