2604.18258v1 Apr 20, 2026 cs.CV

조합적 프롬프트 분해를 통한 장문 텍스트-이미지 생성

Long-Text-to-Image Generation via Compositional Prompt Decomposition

Jen-Yuan Huang
Jen-Yuan Huang
Citations: 8
h-index: 2
Tong Lin
Tong Lin
Citations: 94
h-index: 4
Yilun Du
Yilun Du
Citations: 72
h-index: 2

최신 텍스트-이미지(T2I) 모델은 복잡한 프롬프트로부터 이미지를 생성하는 데 뛰어난 성능을 보이지만, 입력이 상세한 문단일 경우 핵심 내용을 제대로 반영하는 데 어려움을 겪습니다. 이러한 제한은 T2I 모델의 학습 데이터가 주로 간결한 설명을 포함하고 있기 때문입니다. 기존 방법들은 T2I 모델을 긴 프롬프트로 미세 조정하거나, 과도하게 긴 입력을 일반적인 프롬프트 공간으로 투영하여 충실도를 손상시키는 방식으로 이 격차를 해소하려고 시도합니다. 우리는 사전 학습된 T2I 모델이 긴 시퀀스 입력을 처리할 수 있도록 하는 조합적 접근 방식인 프롬프트 굴절을 통한 복잡한 장면 모델링(PRISM)을 제안합니다. PRISM은 가벼운 모듈을 사용하여 긴 프롬프트에서 구성 요소를 추출합니다. T2I 모델은 각 구성 요소에 대해 독립적인 노이즈 예측을 수행하고, 에너지 기반 결합을 사용하여 해당 결과를 단일 디노이징 단계로 병합합니다. 우리는 다양한 모델 아키텍처에서 PRISM을 평가하여 동일한 학습 데이터를 사용하여 미세 조정된 모델과 유사한 성능을 보이는 것을 확인했습니다. 또한, PRISM은 뛰어난 일반화 성능을 보여주며, 어려운 공개 벤치마크에서 500 토큰 이상의 프롬프트에 대해 기준 모델보다 7.4% 더 높은 성능을 달성했습니다.

Original Abstract

While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are descriptive paragraphs. This limitation stems from the prevalence of concise captions that shape their training distributions. Existing methods attempt to bridge this gap by either fine-tuning T2I models on long prompts, which generalizes poorly to longer lengths; or by projecting the oversize inputs into normal-prompt space and compromising fidelity. We propose Prompt Refraction for Intricate Scene Modeling (PRISM), a compositional approach that enables pre-trained T2I models to process long sequence inputs. PRISM uses a lightweight module to extract constituent representations from the long prompts. The T2I model makes independent noise predictions for each component, and their outputs are merged into a single denoising step using energy-based conjunction. We evaluate PRISM across a wide range of model architectures, showing comparable performances to models fine-tuned on the same training data. Furthermore, PRISM demonstrates superior generalization, outperforming baseline models by 7.4% on prompts over 500 tokens in a challenging public benchmark.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!