2605.26955v1 May 26, 2026 cs.CL

JuICE: 문화적 오류 식별을 위한 LLM 평가 벤치마크

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

Sunipa Dev
Sunipa Dev
University of Utah
Citations: 11,043
h-index: 17
V. Prabhakaran
V. Prabhakaran
Citations: 6
h-index: 1
Jiho Jin
Jiho Jin
Citations: 577
h-index: 9
Junho Myung
Junho Myung
Citations: 695
h-index: 11
Juhyun Oh
Juhyun Oh
Citations: 73
h-index: 5
Junyeong Park
Junyeong Park
Citations: 172
h-index: 4
Rifki Afina Putri
Rifki Afina Putri
KAIST
Citations: 429
h-index: 9
Alice Oh
Alice Oh
Citations: 18
h-index: 3

대규모 언어 모델(LLM)이 전 세계 사용자에게 점점 더 많이 배포되면서, 개인적인 의사소통 작성부터 창의적인 아이디어 구상까지 다양한 문화적 맥락에서 일상 업무에 통합되고 있습니다. 이러한 작업은 본질적으로 문화적 특성을 가지며, 문맥적 적절성, 상징적 공감, 그리고 암묵적인 문화적 기대를 필요로 합니다. 원어민들은 이러한 요소들을 직관적으로 활용하며, 따라서 사실적으로는 타당할 수 있지만 현지 독자에게는 명백히 잘못된 답변이 나올 수 있습니다. 기존의 문화 관련 벤치마크는 문장 검증 또는 규범 함축 방법을 통해 문화를 단순한 사실들의 집합으로 취급했으며, LLM을 평가 도구로 사용했지만 이러한 심층적인 문화적 오류를 제대로 파악하는지 여부는 고려하지 않았습니다. 이러한 문제점을 해결하기 위해, 우리는 JuICE(LLM이 문화적 오류를 식별하는 데 사용되는 벤치마크)라는 다국어 데이터 세트를 제안합니다. 이 데이터 세트는 장문의 LLM 응답에서 발견된 7,470개의 문화 및 언어적 오류에 대한 스팬 수준의 주석으로 구성되어 있습니다. 여기에는 미국, 한국, 인도네시아, 방글라데시의 네 개 국가(영어와 각 국가의 주요 언어)에서 수집한 1,050개의 질의-응답 쌍이 포함되어 있습니다. JuICE를 사용하여 분석한 결과, 가장 강력한 LLM 평가 모델조차도 오류 스팬 탐지 작업에서 F1 값이 0.52에 불과하다는 것을 확인했습니다. 또한, LLM 평가 모델은 현지 주민들이 쉽게 식별하는 심층적인 문화적 오류를 지속적으로 놓치고 있습니다. 이러한 결과는 견고한 문화 평가가 표면적인 감지에 머무르지 않고 문화적 의미의 깊이와 맥락성을 고려하는 프레임워크로 발전해야 함을 시사합니다.

Original Abstract

As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural contexts, from drafting personal communications to brainstorming creative ideas. These tasks are inherently cultural: they require contextual appropriateness, symbolic resonance, and tacit cultural expectations that native speakers draw on instinctively, meaning that a response can be factually plausible yet unmistakably wrong to a local reader. Existing cultural benchmarks have treated culture as a flat set of facts via fact verification or norm entailment methods, and have adopted LLM-as-a-Judge without examining whether they can capture such thick cultural errors. To address this gap, we present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both English and their countries' main languages. Using JuICE, we find that even the strongest LLM-judge achieves only an F1 of 0.52 in the erroneous span detection task. Furthermore, LLM-judges consistently miss thick cultural errors that local residents readily identify. Our findings suggest that robust cultural evaluation must move beyond surface-level detection toward frameworks that account for the depth and situatedness of cultural meaning.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!