2604.06863v1 Apr 08, 2026 cs.SI

디지털 피부, 디지털 편향: LLM 및 이모지 임베딩에서 나타나는 톤 기반 편향 분석

Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji Embeddings

Mingchen Li
Mingchen Li
Citations: 33
h-index: 3
Wajdi M Aljedaani
Wajdi M Aljedaani
Citations: 1,493
h-index: 23
Yingjie Liu
Yingjie Liu
Citations: 0
h-index: 0
Navyasri Meka
Navyasri Meka
Citations: 0
h-index: 0
Xuan Lu
Xuan Lu
Citations: 16
h-index: 2
Xinyue Ye
Xinyue Ye
Citations: 45
h-index: 2
Junhua Ding
Junhua Ding
Citations: 49
h-index: 3
Yunhe Feng
Yunhe Feng
Citations: 50
h-index: 4

피부색 이모지는 온라인 커뮤니케이션에서 개인 정체성과 사회적 포용을 증진하는 데 중요한 역할을 합니다. 특히 대규모 언어 모델(LLM)과 같은 AI 모델이 웹 플랫폼 상의 상호작용을 점점 더 많이 중재함에 따라, 이러한 시스템이 피부색 이모지의 표현을 통해 사회적 편향을 지속시키는 위험은 심각한 문제입니다. 본 논문은 두 가지 유형의 모델에 대한 피부색 이모지 표현의 편향을 비교 분석하는 최초의 대규모 연구를 제시합니다. 우리는 전용 이모지 임베딩 모델(emoji2vec, emoji-sw2v)을 현대적인 4가지 LLM(Llama, Gemma, Qwen, Mistral)과 체계적으로 비교 평가했습니다. 분석 결과, LLM은 피부색 수정자에 대한 강력한 지원을 보이는 반면, 널리 사용되는 특수 이모지 모델은 심각한 성능 저하를 보이는 중요한 성능 격차가 있음을 확인했습니다. 더욱 중요하게는, 의미 일관성, 표현 유사성, 감정 극성 및 핵심 편향에 대한 다각적인 조사를 통해 체계적인 불일치가 드러났습니다. 다양한 피부색에 따라 이모지와 관련된 감정의 왜곡 및 일관성 없는 의미가 존재한다는 증거를 발견했으며, 이는 이러한 기반 모델 내에 잠재적인 편향이 있음을 시사합니다. 본 연구 결과는 개발자 및 플랫폼이 이러한 표현적 피해를 감사하고 완화할 필요가 있음을 강조하며, AI가 웹에서 진정한 공정성을 증진하고 사회적 편향을 강화하지 않도록 해야 합니다.

Original Abstract

Skin-toned emojis are crucial for fostering personal identity and social inclusion in online communication. As AI models, particularly Large Language Models (LLMs), increasingly mediate interactions on web platforms, the risk that these systems perpetuate societal biases through their representation of such symbols is a significant concern. This paper presents the first large-scale comparative study of bias in skin-toned emoji representations across two distinct model classes. We systematically evaluate dedicated emoji embedding models (emoji2vec, emoji-sw2v) against four modern LLMs (Llama, Gemma, Qwen, and Mistral). Our analysis first reveals a critical performance gap: while LLMs demonstrate robust support for skin tone modifiers, widely-used specialized emoji models exhibit severe deficiencies. More importantly, a multi-faceted investigation into semantic consistency, representational similarity, sentiment polarity, and core biases uncovers systemic disparities. We find evidence of skewed sentiment and inconsistent meanings associated with emojis across different skin tones, highlighting latent biases within these foundational models. Our findings underscore the urgent need for developers and platforms to audit and mitigate these representational harms, ensuring that AI's role on the web promotes genuine equity rather than reinforcing societal biases.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!