시각 생성에서의 텍스트 조건부 학습의 확장 특성
Scaling Properties of Text Conditioning in Visual Generation
본 연구에서는 시각 생성 과정에서 텍스트 조건을 부여할 때 나타나는 경험적인 확장 특성을 분석합니다. 이러한 특성은 자연어 프롬프트 내 토큰 수에 따라 확산 손실이 변하지 않기 때문에, 그동안 제대로 측정되지 않았습니다. 놀랍게도, 우리는 수렴된 확산 손실이 프롬프트 내 구조화된 언어의 양과 함께 증가하는 것을 발견했습니다. 구조화된 언어를 정량적으로 평가하기 위해, 우리는 두 가지 상호 보완적인 지표를 사용했습니다. 첫째는 내부 정보를 활용한 가능성 측정 지표(GPG)이고, 둘째는 외부 정보를 활용한 속성 측정 지표(ED)입니다. 통제된 학습 실험을 통해, 수렴된 확산 손실은 GPG 값에 따라 거의 선형적으로 감소하며, ED 값에 대해서는 거듭제곱 법칙을 따릅니다. 이러한 확장 특성을 바탕으로, 우리는 이미지에서 파생된 의미론적 및 기하학적 주석을 사용하여 구조화된 프롬프트를 구성함으로써 extit{확산 가능성(diffusability)}을 향상시켰고, 지도 학습, 초기 학습 및 검증기 기반 온폴리시 증류를 통해 프로mp터를 훈련시켜 extit{프롬프트 민감도(promptability)}를 개선했습니다. 결과적으로 개발된 시스템은 거의 모든 합성, 추론 및 세상 지식 평가 기준으로 기존의 공개 모델보다 우수한 성능을 보였으며, 대부분의 평가에서 가장 강력한 비공개 모델과 동등하거나 뛰어넘는 성능을 달성했습니다.
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.