G-IdiomAlign: 어휘 해설(Gloss)을 활용한 다국어 관용구 정렬 벤치마크
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
관용구는 비조성적 특성과 약한 표면 형태 연관성으로 인해 언어 간 번역이 어렵습니다. 따라서 문자 그대로의 매핑은 신뢰할 수 없습니다. 본 논문에서는 Wiktionary의 영어 어휘 해설을 사용하여 각 관용구를 연결하는 어휘 해설 기반 벤치마크인 G-IdiomAlign을 제시합니다. 또한, 재현 가능한 평가를 위해 높은 신뢰도를 가진 참조 정렬 세트를 구축했습니다. G-IdiomAlign은 다음 두 가지 프로토콜을 지원합니다: (1) 오답 유형이 지정된 다중 선택 관용구 동등성 문제로, 오류 원인 분석에 활용됩니다; (2) 어휘 해설 유무를 대비하는 어휘 해설 대조 생성 방식으로, 명시적인 의미 기반 연결 고리가 미치는 영향을 파악합니다. 다양한 LLM을 대상으로 실험한 결과, 문자 그대로 번역에 대한 경향이 주요 실패 요인이었으며, 특히 대상 언어가 자원이 부족한 경우 더욱 두드러졌습니다. 어휘 해설은 임베딩 기반의 의미적 대리(semantic proxy) 하에서 어휘 해설 대조 생성 성능을 꾸준히 향상시키지만, 여전히 미미한 수준이며 이는 추가적인 개선 가능성이 존재함을 시사합니다. Qwen3-8B 모델에 대한 추가 분석 결과, 조건 간 차이가 레이어보다는 어텐션 헤드에 더 집중되어 있으며, 더 나은 어휘 해설 생성 결과는 강력한 어휘 해설 연결과 관련이 있는 것으로 나타났습니다.
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary. We further construct a high-confidence reference alignment set for reproducible evaluation. G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom Equivalence with typed distractors for error attribution; and (2) a Gloss-Contrastive Generation contrasting No-gloss and With-gloss inputs to isolate the effect of an explicit semantic pivot. Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low-resource language. Glosses consistently improve Gloss-Contrastive Generation under an embedding-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space. Subsequent analysis on Qwen3-8B further suggests that cross-condition differences are concentrated more in attention heads than in layers, while better With-gloss generations coincide with stronger gloss anchoring.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.