개방형 작업에서의 언어 모델 창의성 자동 평가
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
대규모 언어 모델(LLM)은 언어 이해, 추론 및 생성 분야에서 놀라운 발전을 이루었으며, 이는 LLM의 잠재적인 창의성에 대한 관심 증가로 이어지고 있습니다. 이러한 잠재력을 실현하기 위해서는 다양한 작업에 걸쳐 창의성을 평가하는 체계적이고 확장 가능한 방법이 필요합니다. 그러나 대부분의 기존 창의성 측정 지표는 특정 작업과 밀접하게 연결되어 있어, 평가 과정에 도메인 특유의 가정을 내포하고 있으며, 이는 확장성과 일반성을 제한합니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 개방형 작업에서 LLM의 창의성을 정량화하는 데 사용될 수 있는 자동화된, 도메인 독립적인 프레임워크를 제시합니다. 우리 접근 방식은 측정 장치를 실제 창의적 작업과 분리하여 확장 가능하고 작업에 독립적인 평가를 가능하게 합니다. 발산적 창의성은 참조가 필요 없는 강력한 새로운 지표 및 다양성 측정 방법인 의미론적 엔트로피를 사용하여 측정하며, 인간 주석, LLM 기반 참신성 판단 및 기준 다양성 측정과 함께 검증되었습니다. 수렴적 창의성은 새롭게 제안된 검색 기반 다중 에이전트 평가 프레임워크를 통해 평가되며, 이는 60% 이상의 효율성을 향상시킨 맥락에 민감한 작업 수행 능력 평가를 제공합니다. 본 연구에서는 문제 해결(MacGyver), 연구 아이디어 생성(HypoGen) 및 창의적 글쓰기(BookMIA) 등 질적으로 구별되는 세 가지 분야에서 다양한 LLM을 사용하여 프레임워크를 검증했습니다. 실험 결과는 우리 프레임워크가 참신성, 다양성 및 작업 수행 능력과 같은 창의성의 핵심 측면을 안정적으로 포착하며, 모델 크기, 온도, 최근성 및 추론과 같은 모델 특성이 창의적 성능에 미치는 영향을 보여줍니다. 본 연구는 자동화된 LLM 창의성 평가를 위한 재현 가능하고 일반화 가능한 표준을 확립하여, 확장 가능한 벤치마킹을 촉진하고 인공지능 분야의 발전을 가속화할 수 있습니다.
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks. However, most existing creativity metrics are tightly coupled to specific tasks, embedding domain assumptions into the evaluation process, and limiting scalability and generality. To address this gap, we introduce an automated, domain-agnostic framework for quantifying LLM creativity across open-ended tasks. Our approach separates the measurement apparatus from the creative task itself, enabling scalable, task-agnostic assessment. Divergent creativity is measured using semantic entropy, a reference-free and robust metric for novelty and diversity, validated against human annotations, LLM-based novelty judgments and baseline diversity measures. Convergent creativity is assessed via a novel retrieval-based multi-agent judge framework that delivers context-sensitive evaluation of task fulfilment with over 60% improved efficiency. We validate our framework in three qualitatively distinct domains: problem-solving (MacGyver), research ideation (HypoGen), and creative writing (BookMIA), using a broad suite of LLMs. Empirical results show that our framework reliably captures key facets of creativity, including novelty, diversity, and task fulfilment, and reveal how model properties, such as size, temperature, recency, and reasoning, impact creative performance. Our work establishes a reproducible and generalizable standard for automated LLM creativity evaluation, paving the way for scalable benchmarking and accelerating progress in creative AI.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.