2607.14962v1 Jul 16, 2026 cs.LG

텍스트-이미지 생성에서 대표성 다양성을 위한 다중 축 Max@K 강화 학습

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

Hiroki Furuta
Hiroki Furuta
Citations: 2,711
h-index: 16
Shohei Taniguchi
Shohei Taniguchi
Citations: 71
h-index: 4
Soichiro Nishimori
Soichiro Nishimori
Citations: 89
h-index: 4
Ku Onoda
Ku Onoda
Citations: 0
h-index: 0
Paavo Parmas
Paavo Parmas
Citations: 247
h-index: 6
Yuta Oshima
Yuta Oshima
Citations: 111
h-index: 6
Yutaka Matsuo
Yutaka Matsuo
Citations: 112
h-index: 4

텍스트-이미지(T2I) 모델은 현실적이고 프롬프트에 부합하는 이미지를 생성할 수 있지만, 동일한 프롬프트로 생성된 이미지들은 종종 시각적으로 뚜렷한 다양한 형태의 작은 부분만 포함합니다. 이는 이미지의 다양성을 제한하며, 특히 인물 중심의 프롬프트에서 인구 통계적 편향을 반영하거나 증폭시킬 수 있습니다. 우리는 이 문제를 미리 정의된 의미론적 범주의 커버리지 문제로 공식화했으며, 이를 '타겟 모드 커버리지'라고 명명했습니다. 그런 다음, 확산 기반 T2I 모델에서의 이러한 커버리지를 개선하기 위한 그룹 기반 강화 학습 목표인 다중 축 Max@K를 제안합니다. 다중 축 Max@K는 샘플 그룹과 각 타겟 범주에 대한 하나의 점수를 입력받아 먼저 각 범주에서 샘플별 최대 점수를 선택하고, 그 후 이러한 범주별 최대값들을 합산합니다. 결과적으로 얻어지는 보상 할당은 특정 샘플이 해당 범주의 그룹 내 최대값을 증가시키는 경우에만 긍정적인 가중치를 부여하여, 서로 다른 샘플들이 서로 다른 범주에 기여할 수 있도록 합니다. 우리는 먼저 합성 데이터셋과 SD3.5-M에서 결정론적인 픽셀 기반 색상 보상을 사용하여 이 보상 할당 메커니즘을 검증합니다. 그런 다음, 동일한 목표를 사용하여 인지된 외형 공정성을 평가했습니다. 세 가지 자동 평가 도구를 사용하여 테스트 프롬프트에 대해 다중 축 Max@K는 기준 모델 대비 Fairness Score를 0.23~0.36만큼 향상시키는 반면, 이미지 품질과 텍스트 정렬을 유지합니다.

Original Abstract

Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for person-centric prompts, can reflect or amplify demographic skew. We formalize this problem as coverage of a predefined set of semantically specified modes, which we call target-mode coverage. We then propose multi-axis max@K, a group-based reinforcement learning objective for improving such coverage in diffusion-based T2I models. Given a group of samples and one score per target category, multi-axis max@K first takes the maximum score across samples for each category and then sums these category-wise maxima. The resulting credit assignment gives a sample positive weight on a category only when it increases that category's group-wise maximum, allowing different samples to contribute to different categories. We first validate the credit-assignment mechanism on a synthetic mixture and on SD3.5-M using deterministic pixel-based color rewards. We then evaluate the same objective on perceived-appearance fairness. Across three automatic evaluators on held-out prompts, multi-axis max@K improves the Fairness Score by 0.23-0.36 relative to the base model, while maintaining image quality and text alignment.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!