ChildEval: 대규모 언어 모델과 어린이의 개성이 만났을 때
ChildEval: When large language models meet children's personalities
대규모 언어 모델(LLM)은 개인 맞춤형 챗봇을 가능하게 하지만, 아동 중심의 개인화에 있어 LLM의 효과는 아직 명확하지 않으며, 특히 아동 특유의 선호도를 체계적으로 평가하는 연구가 부족합니다. 이러한 격차를 해소하기 위해, 우리는 LLM이 장문의 대화에서 아동 중심의 선호도를 추론하고 따르는 능력을 평가하는 벤치마크인 ChildEval을 소개합니다. ChildEval은 3세부터 6세까지의 어린이 인물 프로필 29,000개를 포함하며, 비교적 정적인 배경 정보를 제공합니다. 각 인물은 명시적으로 한 문장으로 표현되거나, 6~10턴의 대화를 통해 암묵적으로 드러나는 아동의 선호도와 연결되어 있으며, 이 선호도는 인물의 특성과 일치하거나, 반대되거나, 독립적일 수 있습니다. 명시적인 선호도와 암묵적인 선호도는 동일한 근본적인 선호도를 반영하지만, 표현 방식은 다릅니다. 이는 선호도 표현의 역동적인 측면을 포착하며, 정적인 인물의 변화를 나타내는 것이 아닙니다. 이 벤치마크는 아동의 일상생활과 발달을 포괄하는 다섯 가지 최상위 범주와 열네 가지 하위 범주로 구성됩니다. 또한, 우리는 오픈 소스 LLM을 체계적으로 평가하기 위한 세분화된 아동 중심 평가 프로토콜을 제안합니다. 실험 결과는 다양한 개인화 표현 방식이 LLM의 응답에 미치는 영향을 보여주며, ChildEval 데이터셋으로 파인튜닝하면 아동 중심 성능을 향상시킬 수 있음을 시사합니다. 저희 코드는 https://github.com/ziyanluo/ChildEval 에서 확인할 수 있습니다.
While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a benchmark for evaluating LLMs' ability to infer and follow child-centered preferences in long-context conversations. ChildEval contains 29K synthesized persona profiles of children aged 3-6, providing relatively static background information. Each persona is associated with a child preference-which may align with, conflict with, or be independent of the persona-expressed either explicitly in a single sentence or implicitly through 6-10 turn dialogues. Explicit and implicit preferences are designed to reflect the same underlying preference but differ in expression, capturing dynamic aspects of preference expression rather than changes in the static persona. The benchmark spans five top-level and fourteen sub-level categories covering children's daily lives and development. We further propose fine-grained, child-centric evaluation protocols to systematically assess open-source LLMs. Experimental results demonstrate how different personalized representations affect LLM responses and suggest that finetuning on ChildEval can enhance child-centered performance. Our code and dataset are available at https://github.com/ziyanluo/ChildEval.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.