2608.01821v1 Aug 03, 2026 cs.CV

DAVET: 디노이징 인지 시각적 증거 경로 할당을 통한 확산 기반 비전-언어 모델

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

Wuyang Zhang
Wuyang Zhang
Citations: 44
h-index: 3
Fan Xu
Fan Xu
Citations: 10
h-index: 2
Xiangwen Xia
Xiangwen Xia
Citations: 0
h-index: 0
Cheng Yan
Cheng Yan
Citations: 122
h-index: 3
Yongkang Zhou
Yongkang Zhou
Citations: 54
h-index: 3

확산 기반 비전-언어 모델(dVLMs)은 마스크된 응답을 반복적으로 노이즈 제거하면서, 각 노이즈 제거 단계에서 시각적 정보를 활용하여 추론 비용을 발생시킵니다. 오토리거시브 디코딩과 달리, 확산 생성 과정에서는 불확실성이 변화함에 따라 전체 응답을 반복적으로 검토합니다. 본 연구에서는 시각적 정보의 필요량이 노이즈 제거 단계에 따라 크게 달라짐을 분석하고, 이에 기반하여 적응적인 할당 전략을 제안합니다. 기존의 추론 가속화 방법은 주로 디코딩 측면에서의 전략이나, 가지치기 및 병합을 통한 시각적 토큰 압축 방식을 사용하지만, 시각적 정보를 확산 과정에서 변화하는 수요를 가진 자원으로 명시적으로 다루지는 않습니다. 따라서 본 연구에서는 노이즈 제거 과정을 고려한 시각적 증거 경로 할당(DAVET)이라는 훈련 불필요한 프레임워크를 제안합니다. DAVET은 단계에 따라 달라지는 상태 정보를 활용하여 시각적 정보를 할당하며, 미리 정의된 시각적 정보 경로를 기반으로 운영 수요를 통해 시각적 정보 예비량을 설정하고, 각 노이즈 제거 단계에서 경로 위험도를 고려하여 할당량을 조절합니다. 제안하는 방법은 단일 시각적 인코딩을 사용하여 계층적인 시각적 정보 뷰를 구성함으로써, 어떤 시점과 얼마나 많은 시각적 정보가 필요한지 그리고 이러한 시각적 정보 뷰가 어떻게 생성되는지를 분리합니다. LLaDA-V 및 LaViDa라는 두 가지 대표적인 dVLMs 모델에 대해 다양한 시각적 이해 벤치마크에서 실험한 결과, DAVET은 평균적으로 1.55배의 속도 향상을 보였으며, 상대적인 성능 저하율은 평균 1.86%로 나타났습니다. 이는 노이즈 제거 과정을 고려한 시각적 정보 할당 전략이 시각적 정보 활용 비용을 줄이는 동시에 생성 품질을 크게 유지할 수 있음을 보여줍니다.

Original Abstract

Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial recurring inference cost. Unlike autoregressive decoding, diffusion generation repeatedly revisits the entire response as uncertainty evolves. Our analysis reveals that visual evidence demand is strongly step-dependent, motivating adaptive allocation across denoising steps. Existing inference acceleration methods operate through decoding-side strategies or visual token compression via pruning and merging, but do not explicitly treat visual evidence as a resource whose demand evolves across the diffusion process. Therefore, we present Denoising-Aware Visual Evidence Trajectory Allocation (DAVET), a training-free framework that allocates visual evidence according to the evolving generation state. Starting from a phase-conditioned evidence trajectory, the proposed allocation policy uses operation demand to set an evidence reserve whose allocation at each denoising step is modulated by trajectory risk. DAVET realizes the resulting budgets through a hierarchy of evidence views constructed from a single visual encoding, separating when and how much evidence is needed from how the evidence views are constructed. Evaluated on two representative dVLMs, LLaDA-V and LaViDa, across multiple visual-understanding benchmarks, DAVET achieves an average speedup of 1.55$\times$ with an average relative performance drop of 1.86\%, showing that denoising-aware visual evidence allocation can reduce visual conditioning cost while largely preserving generation quality.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!