NOLLI: 영어-한국어 성능 격차 진단을 위한 난이도 조정 퍼즐 벤치마크
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
본 논문에서는 한국어 성능 격차가 발생하는 지점을 진단하기 위해 설계된, 절차적으로 생성되는 영어-한국어 퍼즐 벤치마크인 NOLLI를 소개합니다. NOLLI는 15가지 유형의 퍼즐(25개의 과제, 총 7,500개 항목)로 구성되어 있으며, 각 문제는 재현 가능하도록 설계되었으며, 고유한 해답을 가지도록 검증되었고, 결정적으로 점수가 매겨집니다. 우리는 단순히 더 어려운 문제가 크다고 간주하는 대신, 행동적인 기준으로 난이도를 조정하며, 각 생성기를 조정하여 특정 기준 모델이 목표 정확도 범위에 도달하도록 합니다. NOLLI는 세 단계의 설계를 특징으로 하며, 여기에는 직접 번역 문제, 한글 자모(음절 이하 문자)를 활용한 스크립트 변형 문제, 그리고 한국 문화 또는 문법에 기반한 한국어 전용 문제가 포함됩니다. 우리는 15개의 최첨단, 공개 가중치 모델 및 한국에서 개발된 모델을 평가했습니다. 전체 정확도가 3% 이상인 12개 모델 중, 영어-한국어 번역의 정확도는 통계적으로 동등하며 (+/- 10%p) 프레젠테이션 언어 자체만으로는 큰 영향을 미치지 않는 것으로 나타났습니다. 문자 체계를 집중적으로 사용하는 과제에서는 더 큰 격차가 나타났습니다. 예를 들어, 한국어 암호는 영어에 비해 최대 68.7%p 뒤쳐지는 반면, 동일한 자모를 사용한 산술 연산은 체계적인 불이익을 받지 않으며, 자모 조합 정확도는 한국어 암호 정확도를 예측합니다. 이러한 차이는 인과 관계보다는 진단적인 의미를 가지며, 다단계 음절 단위 실행의 어려움과 일관성을 보입니다. 한국어 전용 과제는 규칙 적용 능력 부족(부정적 또는 긍정적 변화)과 모든 12개 모델에서 나타나는 친족 관계 이해 부족을 구분합니다. 마지막으로, 눈에 띄는 크기 지표가 15가지 유형 중 7가지 유형에서 쉬움 단계부터 어려움 단계로 증가하지 않아, 구조적인 크기는 경험적인 난이도의 신뢰할 수 있는 지표가 아님을 보여줍니다.
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.