내러티브 이미지 생성에서의 사회적 편향 연구
Investigating Social Bias in Narrative Image Generation
텍스트-이미지(T2I) 생성 모델은 미디어 콘텐츠 제작 및 교육과 같은 다양한 애플리케이션에 점점 더 많이 활용되고 있으며, 이러한 모델의 결과물이 사회적 편향을 재현할 가능성에 대한 우려가 제기되고 있습니다. 기존 연구에서는 T2I 모델이 사회적 편향을 나타내는 것으로 확인되었지만, 대부분의 평가는 사진 생성 작업에 초점을 맞추고 있습니다. 따라서 이러한 편향이 스토리보드나 만화와 같이 여러 패널로 구성된 내러티브 시각 형식을 통해 어떻게 드러나는지 불분명합니다. 본 연구에서는 텍스트 기반 편향 평가 프레임워크인 BBG를 이미지 생성으로 확장하여, 여섯 개의 T2I 모델에서 사진, 스토리보드 및 만화 생성을 비교했습니다. 결과적으로, 독점 모델은 평균적으로 사진 생성에서 25.9%의 편향된 결과를 생성하며, 스토리보드 생성에서는 9.6%, 만화 생성에서는 18.2%로 편향된 결과가 증가하는 것으로 나타났습니다. 또한, 사진은 주로 미묘한 시각적 단서를 통해 편향을 드러내는 반면, 스토리보드와 만화는 이벤트 순서, 캐릭터 배치, 서사 전개 및 텍스트 요소 등을 통해 이러한 편향을 더 명확하게 드러내는 것을 확인했습니다. 이러한 결과는 사진 생성에서는 덜 눈에 띄는 편향이 내러티브 시각 형식에서 표면화될 수 있음을 보여주며, 사진 생성 외에도 다양한 시각 형식을 사용하여 T2I 시스템을 평가하는 것의 중요성을 강조합니다.
Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.