타일 기반 프롬프트: 이미지 및 비디오 초해상도에서 프롬프트 부족 문제를 해결하는 방법
Tiled Prompts: Overcoming Prompt Underspecification in Image and Video Super-Resolution
텍스트 기반 확산 모델은 프롬프트를 의미론적 사전 정보로 활용하여 이미지 및 비디오 초해상도 성능을 향상시켰지만, 최신 초해상도 파이프라인은 일반적으로 높은 해상도로 확장하기 위해 잠재 이미지 타일링을 사용하는데, 이때 단일의 전역적인 캡션은 프롬프트 부족 문제를 야기합니다. 거친 전역 프롬프트는 종종 지역적인 세부 사항을 놓치고(프롬프트 희소성), 지역적으로 관련 없는 지침을 제공하며(프롬프트 오도), 이는 분류기 자유 지침에 의해 증폭될 수 있습니다. 본 논문에서는 이미지 및 비디오 초해상도를 위한 통합 프레임워크인 타일 기반 프롬프트를 제안합니다. 이 방법은 각 잠재 이미지 타일에 대해 타일별 프롬프트를 생성하고, 지역적으로 텍스트 기반 조건부 사후 분포를 사용하여 초해상도를 수행합니다. 이를 통해 풍부한 정보를 제공하여 프롬프트 부족 문제를 최소한의 오버헤드로 해결합니다. 고해상도 실세계 이미지 및 비디오에 대한 실험 결과, 제안하는 방법은 전역 프롬프트 기반의 기존 방법보다 시각적 품질과 텍스트 일관성 측면에서 꾸준히 향상된 성능을 보이며, 환각 현상 및 타일 수준의 결함을 줄이는 것을 확인했습니다.
Text-conditioned diffusion models have advanced image and video super-resolution by using prompts as semantic priors, but modern super-resolution pipelines typically rely on latent tiling to scale to high resolutions, where a single global caption causes prompt underspecification. A coarse global prompt often misses localized details (prompt sparsity) and provides locally irrelevant guidance (prompt misguidance) that can be amplified by classifier-free guidance. We propose Tiled Prompts, a unified framework for image and video super-resolution that generates a tile-specific prompt for each latent tile and performs super-resolution under locally text-conditioned posteriors, providing high-information guidance that resolves prompt underspecification with minimal overhead. Experiments on high resolution real-world images and videos show consistent gains in perceptual quality and text alignment, while reducing hallucinations and tile-level artifacts relative to global-prompt baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.