더 빠르지만 다름: 가속화된 멀티모달 디퓨전 언어 모델에서 발생하는 콘텐츠 변화 진단 및 제어
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
학습 과정 없이 가속화를 통해 구현 가능성이 높아진 디퓨전 기반의 멀티모달 대규모 언어 모델(dMLLM)은 생성되는 콘텐츠를 서서히 변경할 수 있습니다. 본 연구에서는 300개의 실제 이미지를 사용하여, Fast-dLLM의 결과와 동일 모델의 가속화되지 않은 결과 간의 일관성 문제를 분석합니다. 긴 형식 설정에서 유도된 경미한 병렬 처리(단계당 1.05~1.25개의 토큰) 하에서, 신뢰도 임계값 조정은 디코딩 동작을 변경하지만 기본 수준의 일치성은 유지됩니다. 상태 재시작 실험 및 이미지 교환 개입 결과, 오래된 시각적 정보와 생성된 텍스트 상태가 콘텐츠 변화의 원인이 되는 것으로 확인되었습니다. 테스트된 Fast-dLLM 구현에서 KV 캐시 재생성 간격을 줄이면 속도와 일치성 사이의 단조적인 관계를 형성하며, 측정된 1.3배의 속도 향상 시 거의 완벽한 일치성을 얻을 수 있습니다. 초기 진단 결과는 dLLM-Cache 및 LaViDa에서도 나타나지만, dLLM-Cache는 캐시를 모두 조인해야만 일치성이 회복되어 속도 이점을 잃게 됩니다. 독립적인 프롬프트와 이미지를 사용하여 임계값에 대한 민감성 부족과 재생성 복구 현상이 재현됩니다. 표적 감사 결과, 50쌍의 낮은 일치성을 보이는 데이터 중 절반에서 실제 콘텐츠 대체가 발생하는 것으로 확인되었습니다. 별도의 이중 검토자 평가에서는 가속화된 모델과 기준 모델 간의 사실 오류 차이가 평균 0.00(95% 신뢰 구간 [-0.17, +0.17])으로 나타났습니다. 이는 해당 샘플에서 차이를 감지하지 못했지만, 사실적 동등성을 입증하는 것은 아닙니다. 마지막으로, 테스트된 모든 적응형 또는 부드러운 재생성 방식이 동일한 계산량에서 고정 간격 방식을 능가하지 못했습니다. 본 연구의 기여는 쌍을 이루는 진단 방법과 구현 범위 내에서의 일관성 제어 기술이며, 정확도나 안전성을 보장하는 것은 아닙니다.
Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing Fast-dLLM outputs with the same model's unaccelerated outputs. Across the mild parallelism induced in our long-form setting (1.05--1.25 committed tokens per step), confidence-threshold tuning changes decoding behavior but not baseline agreement. State-refresh ablations and an image-swap intervention instead identify stale visual and generated-text states as contributors to drift. For the tested Fast-dLLM implementation, shortening the KV-cache refresh interval yields a monotonic speed--agreement frontier and near-exact agreement at a measured 1.3x speedup. The initial diagnosis also appears with dLLM-Cache and LaViDa, although dLLM-Cache recovers agreement only after both caches are tightened, which removes its speed advantage. Independent prompts and images reproduce the threshold-insensitivity and refresh recovery. A targeted audit finds genuine content substitution in half of 50 low-agreement pairs. In a separate blinded two-annotator evaluation, the pooled accelerated-minus-baseline factual-error difference is 0.00 (95% CI [-0.17,+0.17]); this sample detects no difference but does not establish factual equivalence. Finally, none of the tested adaptive or smoothed-refresh variants beats the fixed interval at matched compute. Our contribution is a paired diagnostic and an implementation-scoped consistency control, not an accuracy or safety guarantee.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.