LumaGuide: 디퓨전 모델에서 학습 없이 HDR 이미지를 생성하기 위한 분포 조정 방법
LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models
사전에 학습된 디퓨전 모델은 사실적인 이미지를 생성하지만, 학습 데이터의 통계적 편향으로 인해 고다이나믹 레인지(HDR) 콘텐츠를 생성하는 데 제약이 있습니다. 본 연구에서는 디퓨전 모델에서 학습 없이 분포를 조정할 수 있는 프레임워크인 LumaGuide를 소개합니다. LumaGuide는 모델 파라미터를 수정하는 대신, 미분 가능한 에너지 기반 가이드랜스를 통해 샘플링 과정을 제어하여 목표 특징 분포와 일치하도록 합니다. 우리는 이 프레임워크를 HDR 이미지 생성에 적용하여 PQ 공간에서 시각적으로 균일한 휘도 분포를 제어합니다. 실험 결과는 휘도 히스토그램을 정렬하는 것만으로도 HDR과 일관된 결과를 얻을 수 있으며, 이는 응집력 있는 하이라이트와 보존된 그림자 디테일을 포함합니다. 또한 의미론적 정확성을 유지합니다. LumaGuide는 데이터 기반 프리셋, 참조 이미지 또는 텍스트 기반 예측기를 통해 목표 분포를 유연하게 지정할 수 있도록 하며, 시간 일관성 제약 조건과 함께 비디오 생성에도 자연스럽게 적용될 수 있습니다. 더욱 광범위하게, 본 연구는 디퓨전 모델을 재학습하지 않고 샘플링 시에 출력 분포를 직접적으로 조정함으로써 제어 가능한 이미지 생성이 가능하다는 것을 보여줍니다.
Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to produce high dynamic range (HDR) content. In this work, we introduce LumaGuide, a training-free framework for distribution shaping in diffusion models. Instead of modifying model parameters, LumaGuide steers the sampling process to match target feature distributions via differentiable energy-based guidance. We instantiate this framework for HDR generation by controlling luminance distributions in perceptually uniform PQ space. Our results show that aligning luminance histograms is sufficient to induce HDR-consistent behavior, including coherent highlights and preserved shadow detail, while maintaining semantic fidelity. Beyond HDR, LumaGuide enables flexible specification of target distributions through data-driven presets, reference images, or text-driven predictors, and extends naturally to video generation with temporal consistency constraints. More broadly, our work demonstrates that controllable generation can be achieved by directly shaping output distributions at sampling time, without retraining diffusion models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.