Arena-T2I Hard: 종속성 기반 체크리스트를 활용한 충실도 평가 및 개선
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
텍스트-이미지(T2I) 모델의 실제 유용성에 있어, 생성된 이미지가 프롬프트와 얼마나 정확하게 일치하는가 하는 '충실도'는 점점 더 중요해지고 있습니다. 기존의 충실도 벤치마크는 단순한 기본 명령어를 기반으로 하며, 최상위 시스템은 이미 거의 완벽에 가까운 점수를 달성합니다. T2I 모델이 창작 워크플로우에 진입하면서, 사용자들은 복잡한 공간 관계, 스타일 제약 및 정교한 텍스트 표현을 결합한 다면적인 요청을 내놓습니다. 이러한 상황에서 단일 이진 VLM-판별 점수는 모델이 충족하지 못하는 구체적인 제약을 제대로 반영하지 못합니다. 본 논문에서는 실제 T2I 로그에서 추출된 310개의 프롬프트로 구성된 스트레스 테스트 벤치마크인 Arena-T2I Hard를 소개합니다. 이 벤치마크는 각 프롬프트를 약 30개의 분해된 예/아니오 제약 조건으로 구성하며, 텍스트 표현을 포함한 6가지 범주를 포괄합니다. 우리가 평가한 가장 강력한 비공개 시스템은 0.855의 성능을 보였으며, 11개의 시스템 간에 최대 33%p의 성능 격차를 보여 상당한 차별성을 나타냅니다. 또한, 공개 T2I 순위가 충실도를 예측하지 못한다는 점은 전체적인 Bradley-Terry (BT) 선호도 점수가 미세한 프롬프트 준수보다 심미성을 우선시한다는 것을 확인시켜 줍니다. 우리는 각 프롬프트를 예/아니오 질문으로 구성된 방향성 비순환 그래프(DAG)로 분해하고, 실패한 부모 노드의 자식 노드를 제거하여 충실도를 제약 조건별 학습 신호로 변환하는 종속성 기반 체크리스트 보상을 제안합니다. 그룹 분리 정규화(GDPO)를 통해 각 보상을 해당 rollout 그룹 내에서 표준화하여 어느 하나도 붕괴되지 않도록 하는 BT 심미적 보상과 결합함으로써, SD3.5-Medium 및 FLUX.1-dev 환경에서 MMRB2 쌍 비교 테스트 결과 모든 단일 보상, 단순 가중 합 또는 4개의 보상을 사용하는 BT 앙상블 기준보다 더 나은 충실도-심미성 균형을 달성했습니다.
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users issue multi-faceted requests combining intricate spatial relationships, stylistic constraints, and complex text rendering. In this setting, a single binary VLM-judge score no longer captures which specific constraints the model fails to satisfy. We introduce Arena-T2I Hard, a 310-prompt stress benchmark drawn from real arena T2I logs, with approximately 30 decomposed yes/no constraints per prompt spanning six categories, including text rendering. The strongest closed-source system we evaluate reaches 0.855 with a 33~pp performance gap across 11 systems, demonstrating substantial discriminative power. Moreover, high public-arena rankings fail to predict faithfulness, confirming that holistic Bradley-Terry (BT) preference scores prioritize aesthetics over fine-grained prompt adherence. We propose a dependency-aware checklist reward that decomposes each prompt into a DAG of yes/no questions and zeroes descendants of failed parents, turning faithfulness into a per-constraint training signal. Combined with a BT aesthetic reward via group-decoupled normalization (GDPO), which standardizes each reward within its rollout group so neither collapses, the recipe attains a strictly better faithfulness-aesthetics trade-off on SD3.5-Medium and FLUX.1-dev under MMRB2 pairwise comparisons than every single-reward, naive weighted-sum, or 4-reward BT-ensemble baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.