2607.18724v1 Jul 21, 2026 cs.AI

모든 문제를 해결하는 하나의 수정인가? 텍스트-이미지 프롬프트 최적화를 위한 유형 인식 기반 수리 할당

One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization

Xiaoyu Ma
Xiaoyu Ma
Citations: 15
h-index: 2
Haoyue Liu
Haoyue Liu
Citations: 3
h-index: 1
Xiaoying Tang
Xiaoying Tang
Citations: 492
h-index: 9
Yeheng Chen
Yeheng Chen
Citations: 15
h-index: 2
Shuguang Cui
Shuguang Cui
Citations: 0
h-index: 0

텍스트-이미지(T2I) 생성기는 종종 사용자의 지시를 충실히 따르지 못하여 잘못된 개수, 뒤바뀐 속성, 모호한 관계 및 읽을 수 없는 텍스트와 같은 오류를 발생시키는 경우가 많습니다. 프롬프트 최적화는 이러한 실패를 해결하기 위해 사용자 프롬프트를 재작성하는 방식으로, 생성기 재학습이 필요 없으며 유망한 결과를 보여주었습니다. 그러나 기존 최적화 도구들은 다양한 유형의 오류들을 하나의 통일된 프롬프트 확장으로 처리합니다. 이는 각 오류에 적합한 다른 수리 언어가 필요한 경우 비효율적입니다. 본 연구에서는 의미론적 프롬프트 최적화를 원자적인 수리 할당 문제로 정의합니다. 즉, 각 실패한 명제는 유형에 따라 조건화된 수리 연산자로 라우팅되고, 결과적으로 생성되는 지역 제약 조건은 하나의 실행 가능한 프롬프트로 컴파일됩니다. 이러한 구조를 훈련이 필요 없는 유형 인식 기반 수리 할당(TARA) 프레임워크에서 구현했습니다. TARA는 진단, 할당, 컴파일 단계를 분리하고, 의미론적 오류 회귀를 방지하는 수리 게이트라는 수락/반환 제어기를 사용하여 정확히 하나의 지정된 수리를 수행합니다. DSG 및 TIFA 데이터셋을 사용한 광범위한 실험 결과, TARA는 네 가지 고정된 생성기 모델에서 모두 가장 높은 의미론적 정확도를 달성했으며, VisualPrompter에 비해 각각 5.6점과 2.6점을 향상시켰습니다. 또한 이미지 품질을 유지하면서 동일한 환경 설정에서 프롬프트당 16.0초로 가장 빠른 실행 속도를 보였습니다 (VisualPrompter는 20.0초).

Original Abstract

Text-to-image (T2I) generators often fail to follow their prompts faithfully, producing wrong counts, swapped attributes, ambiguous relations, and illegible text. Prompt optimization repairs such failures by rewriting the user prompt, requiring no generator retraining, and has yielded promising results. However, existing optimizers absorb heterogeneous failures into one uniform prompt expansion, even though each calls for different repair language. We formulate semantic prompt optimization as atomic repair allocation: each failed proposition is routed to a type-conditioned repair operator before the resulting local constraints are compiled into one executable prompt. We instantiate this formulation in the training-free Type-Aware Repair Allocation (TARA) framework, which separates diagnosis, allocation, compilation, and a semantic repair gate, an accept-or-revert controller over exactly one prescribed repair that prevents semantic regressions. Extensive experiments on DSG and TIFA across four frozen generators demonstrate that TARA achieves the best semantic accuracy in all eight benchmark-generator cells, improving over VisualPrompter by 5.6 and 2.6 points on DSG and TIFA, respectively, while maintaining image quality and running fastest in our matched local setting at 16.0 seconds versus 20.0 seconds per prompt.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!