2603.14294v1 Mar 15, 2026 cs.CV

확산 잡음에서 찾은 물리학적 특징

Seeking Physics in Diffusion Noise

Chujun Tang
Chujun Tang
Citations: 0
h-index: 0
Lei Zhong
Lei Zhong
Citations: 346
h-index: 4
Fangqiang Ding
Fangqiang Ding
Citations: 21
h-index: 2

비디오 확산 모델이 물리적 타당성을 예측하는 정보를 포함하고 있는가? 우리는 사전 훈련된 Diffusion Transformer (DiT) 모델의 중간 단계 노이즈 제거 표현을 분석하여, 물리적으로 타당한 비디오와 타당하지 않은 비디오가 노이즈 수준에 따라 중간 레이어의 특징 공간에서 부분적으로 분리될 수 있음을 확인했습니다. 이러한 분리 현상은 시각적 품질이나 생성 모델의 특성으로 완전히 설명될 수 없으며, 이는 동결된 DiT 특징 내에 복구 가능한 물리학 관련 단서가 존재한다는 것을 시사합니다. 이러한 관찰을 바탕으로, 우리는 '점진적 경로 선택(progressive trajectory selection)'이라는 추론 시간 전략을 도입합니다. 이 전략은 사전 훈련된 동결 특징을 사용하여 경량 물리 검증기를 통해 몇 개의 중간 체크포인트에서 병렬 노이즈 제거 경로를 평가하고, 낮은 점수를 받는 후보를 초기 단계에서 제거합니다. PhyGenBench에 대한 광범위한 실험 결과, 우리 방법은 물리적 일관성을 향상시키면서 추론 비용을 줄이며, 현저히 적은 노이즈 제거 단계를 사용하여 Best-of-K 샘플링과 유사한 결과를 달성하는 것을 보여줍니다.

Original Abstract

Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of a pretrained Diffusion Transformer (DiT) and find that physically plausible and implausible videos are partially separable in mid-layer feature space across noise levels. This separability cannot be fully attributed to visual quality or generator identity, suggesting recoverable physics-related cues in frozen DiT features. Leveraging this observation, we introduce progressive trajectory selection, an inference-time strategy that scores parallel denoising trajectories at a few intermediate checkpoints using a lightweight physics verifier trained on frozen features, and prunes low-scoring candidates early. Extensive experiments on PhyGenBench demonstrate that our method improves physical consistency while reducing inference cost, achieving comparable results to Best-of-K sampling with substantially fewer denoising steps.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!