2606.31711v1 Jun 30, 2026 cs.AI

Arena-T2I Hard: 종속성 기반 체크리스트를 활용한 충실도 평가 및 개선

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

I-Hung Hsu
I-Hung Hsu
Citations: 261
h-index: 6
Wei-Lin Chiang
Wei-Lin Chiang
Citations: 14,515
h-index: 22
Ion Stoica
Ion Stoica
Citations: 28,078
h-index: 41
Yuanhao Ban
Yuanhao Ban
Citations: 232
h-index: 5
Yunqi Hong
Yunqi Hong
Citations: 15
h-index: 2
Sohyun An
Sohyun An
Citations: 65
h-index: 4
Tong Xie
Tong Xie
Citations: 37
h-index: 3
Cho-Jui Hsieh
Cho-Jui Hsieh
Citations: 8
h-index: 1
Evan Frick
Evan Frick
Citations: 570
h-index: 4

텍스트-이미지(T2I) 모델의 실제 유용성에 있어, 생성된 이미지가 프롬프트와 얼마나 정확하게 일치하는가 하는 '충실도'는 점점 더 중요해지고 있습니다. 기존의 충실도 벤치마크는 단순한 기본 명령어를 기반으로 하며, 최상위 시스템은 이미 거의 완벽에 가까운 점수를 달성합니다. T2I 모델이 창작 워크플로우에 진입하면서, 사용자들은 복잡한 공간 관계, 스타일 제약 및 정교한 텍스트 표현을 결합한 다면적인 요청을 내놓습니다. 이러한 상황에서 단일 이진 VLM-판별 점수는 모델이 충족하지 못하는 구체적인 제약을 제대로 반영하지 못합니다. 본 논문에서는 실제 T2I 로그에서 추출된 310개의 프롬프트로 구성된 스트레스 테스트 벤치마크인 Arena-T2I Hard를 소개합니다. 이 벤치마크는 각 프롬프트를 약 30개의 분해된 예/아니오 제약 조건으로 구성하며, 텍스트 표현을 포함한 6가지 범주를 포괄합니다. 우리가 평가한 가장 강력한 비공개 시스템은 0.855의 성능을 보였으며, 11개의 시스템 간에 최대 33%p의 성능 격차를 보여 상당한 차별성을 나타냅니다. 또한, 공개 T2I 순위가 충실도를 예측하지 못한다는 점은 전체적인 Bradley-Terry (BT) 선호도 점수가 미세한 프롬프트 준수보다 심미성을 우선시한다는 것을 확인시켜 줍니다. 우리는 각 프롬프트를 예/아니오 질문으로 구성된 방향성 비순환 그래프(DAG)로 분해하고, 실패한 부모 노드의 자식 노드를 제거하여 충실도를 제약 조건별 학습 신호로 변환하는 종속성 기반 체크리스트 보상을 제안합니다. 그룹 분리 정규화(GDPO)를 통해 각 보상을 해당 rollout 그룹 내에서 표준화하여 어느 하나도 붕괴되지 않도록 하는 BT 심미적 보상과 결합함으로써, SD3.5-Medium 및 FLUX.1-dev 환경에서 MMRB2 쌍 비교 테스트 결과 모든 단일 보상, 단순 가중 합 또는 4개의 보상을 사용하는 BT 앙상블 기준보다 더 나은 충실도-심미성 균형을 달성했습니다.

Original Abstract

Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users issue multi-faceted requests combining intricate spatial relationships, stylistic constraints, and complex text rendering. In this setting, a single binary VLM-judge score no longer captures which specific constraints the model fails to satisfy. We introduce Arena-T2I Hard, a 310-prompt stress benchmark drawn from real arena T2I logs, with approximately 30 decomposed yes/no constraints per prompt spanning six categories, including text rendering. The strongest closed-source system we evaluate reaches 0.855 with a 33~pp performance gap across 11 systems, demonstrating substantial discriminative power. Moreover, high public-arena rankings fail to predict faithfulness, confirming that holistic Bradley-Terry (BT) preference scores prioritize aesthetics over fine-grained prompt adherence. We propose a dependency-aware checklist reward that decomposes each prompt into a DAG of yes/no questions and zeroes descendants of failed parents, turning faithfulness into a per-constraint training signal. Combined with a BT aesthetic reward via group-decoupled normalization (GDPO), which standardizes each reward within its rollout group so neither collapses, the recipe attains a strictly better faithfulness-aesthetics trade-off on SD3.5-Medium and FLUX.1-dev under MMRB2 pairwise comparisons than every single-reward, naive weighted-sum, or 4-reward BT-ensemble baseline.

0 Citations
0 Influential
20.5 Altmetric
102.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!