2607.29079v1 Jul 31, 2026 cs.CL

더 빠르지만 다름: 가속화된 멀티모달 디퓨전 언어 모델에서 발생하는 콘텐츠 변화 진단 및 제어

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

Yang Shu
Yang Shu
Citations: 66
h-index: 3
Yao Dou
Yao Dou
Citations: 472
h-index: 10

학습 과정 없이 가속화를 통해 구현 가능성이 높아진 디퓨전 기반의 멀티모달 대규모 언어 모델(dMLLM)은 생성되는 콘텐츠를 서서히 변경할 수 있습니다. 본 연구에서는 300개의 실제 이미지를 사용하여, Fast-dLLM의 결과와 동일 모델의 가속화되지 않은 결과 간의 일관성 문제를 분석합니다. 긴 형식 설정에서 유도된 경미한 병렬 처리(단계당 1.05~1.25개의 토큰) 하에서, 신뢰도 임계값 조정은 디코딩 동작을 변경하지만 기본 수준의 일치성은 유지됩니다. 상태 재시작 실험 및 이미지 교환 개입 결과, 오래된 시각적 정보와 생성된 텍스트 상태가 콘텐츠 변화의 원인이 되는 것으로 확인되었습니다. 테스트된 Fast-dLLM 구현에서 KV 캐시 재생성 간격을 줄이면 속도와 일치성 사이의 단조적인 관계를 형성하며, 측정된 1.3배의 속도 향상 시 거의 완벽한 일치성을 얻을 수 있습니다. 초기 진단 결과는 dLLM-Cache 및 LaViDa에서도 나타나지만, dLLM-Cache는 캐시를 모두 조인해야만 일치성이 회복되어 속도 이점을 잃게 됩니다. 독립적인 프롬프트와 이미지를 사용하여 임계값에 대한 민감성 부족과 재생성 복구 현상이 재현됩니다. 표적 감사 결과, 50쌍의 낮은 일치성을 보이는 데이터 중 절반에서 실제 콘텐츠 대체가 발생하는 것으로 확인되었습니다. 별도의 이중 검토자 평가에서는 가속화된 모델과 기준 모델 간의 사실 오류 차이가 평균 0.00(95% 신뢰 구간 [-0.17, +0.17])으로 나타났습니다. 이는 해당 샘플에서 차이를 감지하지 못했지만, 사실적 동등성을 입증하는 것은 아닙니다. 마지막으로, 테스트된 모든 적응형 또는 부드러운 재생성 방식이 동일한 계산량에서 고정 간격 방식을 능가하지 못했습니다. 본 연구의 기여는 쌍을 이루는 진단 방법과 구현 범위 내에서의 일관성 제어 기술이며, 정확도나 안전성을 보장하는 것은 아닙니다.

Original Abstract

Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing Fast-dLLM outputs with the same model's unaccelerated outputs. Across the mild parallelism induced in our long-form setting (1.05--1.25 committed tokens per step), confidence-threshold tuning changes decoding behavior but not baseline agreement. State-refresh ablations and an image-swap intervention instead identify stale visual and generated-text states as contributors to drift. For the tested Fast-dLLM implementation, shortening the KV-cache refresh interval yields a monotonic speed--agreement frontier and near-exact agreement at a measured 1.3x speedup. The initial diagnosis also appears with dLLM-Cache and LaViDa, although dLLM-Cache recovers agreement only after both caches are tightened, which removes its speed advantage. Independent prompts and images reproduce the threshold-insensitivity and refresh recovery. A targeted audit finds genuine content substitution in half of 50 low-agreement pairs. In a separate blinded two-annotator evaluation, the pooled accelerated-minus-baseline factual-error difference is 0.00 (95% CI [-0.17,+0.17]); this sample detects no difference but does not establish factual equivalence. Finally, none of the tested adaptive or smoothed-refresh variants beats the fixed interval at matched compute. Our contribution is a paired diagnostic and an implementation-scoped consistency control, not an accuracy or safety guarantee.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!