2606.24849v1 Jun 23, 2026 cs.CV

IV-CoT: 구조 인식 텍스트-이미지 생성을 위한 암시적 시각적 사고 체인

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Zixuan Li
Zixuan Li
Citations: 63
h-index: 3
Haokun Lin
Haokun Lin
Citations: 38
h-index: 4
Yicheng Xiao
Yicheng Xiao
Citations: 186
h-index: 7
Zhiwei Li
Zhiwei Li
Citations: 14
h-index: 2
Xinyang Song
Xinyang Song
Citations: 2
h-index: 1
Yong He
Yong He
Citations: 0
h-index: 0
Heng Yao
Heng Yao
Citations: 0
h-index: 0
Ke Ding
Ke Ding
Citations: 37
h-index: 3
Chao-Chan Yu
Chao-Chan Yu
Citations: 1
h-index: 1
Chuanhan Yuan
Chuanhan Yuan
Citations: 1
h-index: 1
Qi Li
Qi Li
Citations: 12
h-index: 2
Zhenan Sun
Zhenan Sun
Citations: 10
h-index: 2
Zelong Zheng
Zelong Zheng
Citations: 0
h-index: 0

통합 다중 모드 대규모 언어 모델(MLLM)은 뛰어난 텍스트-이미지 생성 품질을 달성했지만, 여전히 객체 수, 공간 관계, 속성 연결 및 개략적인 레이아웃과 같은 구조 정보를 정확하게 반영하는 데 어려움을 겪고 있습니다. 이러한 한계는 부분적으로 단일 조건부 스트림 내에서 구조 계획과 시각적 표현이 복잡하게 얽혀 있기 때문이라고 볼 수 있습니다. 이 문제를 해결하기 위해, 우리는 질의 기반 이미지 생성을 위한 잠재적인 시각적 추론 프레임워크인 Implicit Visual Chain-of-Thought (IV-CoT)를 제안합니다. IV-CoT는 시각적 조건을 구조-의미 연결 방식으로 분해하여, 구조 관련 질의가 먼저 잠재적인 시각 계획을 형성하고, 그 후에 의미 관련 질의가 이 계획에 따라 시각적 표현을 생성하도록 합니다. 구조 관련 질의를 안내하기 위해, 우리는 학습 단계에서만 사용되는 스케치 지도(sketch supervision) 방법을 도입합니다. 이를 통해 질의는 추론 과정에서 스케치를 추출하거나 중간 디코딩 과정을 거치지 않고도 스케치로부터 구조 정보를 학습하도록 유도합니다. IV-CoT는 단일 순방향 패스 내에서 암시적인 사고 체인 추론을 수행하며, GenEval 및 T2I-CompBench에서 뛰어난 성능을 보입니다. 시각화 및 분석 결과, 학습된 구조 관련 및 의미 관련 질의가 구조 인식 생성 과정에서 상호 보완적인 역할을 한다는 것을 보여줍니다.

Original Abstract

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object counts, spatial relations, attribute bindings, and coarse layouts must be preserved. We attribute this limitation in part to the entanglement of structural planning and appearance rendering within a single conditioning stream. To address this issue, we propose Implicit Visual Chain-of-Thought (IV-CoT), a latent visual reasoning framework for query-conditioned image generation. IV-CoT decomposes the visual conditioning queries into a structural-to-semantic cascade, where structural queries first form a latent visual plan and semantic queries then render appearance conditioned on this plan. To guide the structural queries, we introduce training-only sketch supervision, which encourages them to capture structure from sketches without requiring sketch extraction or intermediate decoding at inference time. IV-CoT performs implicit CoT reasoning in a single forward pass and achieves superior results on GenEval and T2I-CompBench. Visualizations and analyses demonstrate that the learned structural and semantic queries play complementary roles in structure-aware generation.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!