2606.31026v1 Jun 30, 2026 cs.LG

OTCache: 디퓨전 모델에서 기하학적 특성을 고려한 캐싱을 위한 최적 수송 기반 기술

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models

Shiguo Lian
Shiguo Lian
Citations: 309
h-index: 6
Huanlin Gao
Huanlin Gao
Citations: 29
h-index: 3
Fuyuan Shi
Fuyuan Shi
Citations: 49
h-index: 5
Yantao Li
Yantao Li
Citations: 19
h-index: 2
Qiang Hui
Qiang Hui
Citations: 4
h-index: 1
Yuren You
Yuren You
Citations: 3
h-index: 1
Ting Lu
Ting Lu
Citations: 4
h-index: 1
Chao Tan
Chao Tan
Citations: 55
h-index: 3
Shaoan Zhao
Shaoan Zhao
Citations: 5
h-index: 1
Fang Zhao
Fang Zhao
Citations: 14
h-index: 3
Kai Wang
Kai Wang
Citations: 252
h-index: 5

본 논문에서는 학습이 필요 없는 프레임워크인 OTCache를 제안합니다. OTCache는 캐시 스케줄 예측을 통해 디퓨전 샘플링 속도를 향상시키는 데 사용됩니다. 기존의 그래프 기반 캐싱 방법은 최단 경로를 최적화하여 중복 계산을 줄이지만, 종종 문제가 되는 가산 독립 가정에 의존합니다. 이러한 문제를 해결하기 위해 OTCache는 추론 예산을 기준으로 정책 공간에서 부드러운 변화로 캐시 스케줄을 모델링하며, 이는 최적 수송(OT) 개념에서 영감을 받았습니다. 이 프레임워크는 세 단계로 구성됩니다: (1) 보수적인 예산 하에서 그래프 기반 캐싱 방법을 사용하여 고정밀 참조 스케줄을 얻습니다; (2) Optuna 최적화를 통해 엔드-투-엔드 지각 목표를 사용하여 극도로 낮은 예산 환경에서 간단한 앵커 검색을 수행합니다; (3) 연속적인 변형 표현을 사용하여 참조 및 앵커 정책 간의 분위수 보간을 통해 대상 예산을 위한 스케줄을 예측합니다. FLUX.1 [dev], Qwen-Image, HunyuanVideo에 대한 실험 결과, OTCache는 각각 4.5배, 4.7배, 3.66배의 속도 향상을 달성했으며, 최첨단 캐싱 기준 성능보다 생성 품질이 꾸준히 개선되었습니다. 본 연구는 최적 수송 기반 스케줄 모델링을 통해 디퓨전 모델 가속화에 대한 새로운 관점을 제시합니다. 코드: https://github.com/UnicomAI/OTCache

Original Abstract

We propose OTCache, a training-free framework for accelerating diffusion sampling via caching schedule prediction. Existing graph-based caching methods reduce redundant computation by optimizing shortest-path objectives, but rely on an additive independence assumption, which often breaks down in the low NFE regime. To address this issue, OTCache models caching schedules across inference budgets as a smooth evolution in policy space, inspired by Optimal Transport (OT). The framework consists of three stages: (1) obtaining a high-fidelity \textbf{reference schedule} using a graph-based caching method under a conservative budget; (2) performing a lightweight anchor search under an extreme low-budget setting via Optuna optimization with an end-to-end perceptual objective; and (3) predicting schedules for target budgets via quantile interpolation between the reference and anchor policies using continuous warping representations. Experiments on FLUX.1 [dev], Qwen-Image, and HunyuanVideo show that OTCache achieves 4.5x, 4.7x, and 3.66x acceleration, respectively, while consistently improving generation fidelity over state-of-the-art caching baselines. This work provides a new perspective on accelerating diffusion models through Optimal-Transport-inspired schedule modeling. Code:https://github.com/UnicomAI/OTCache

0 Citations
0 Influential
29.931471805599 Altmetric
0.0 Score
Original PDF
3

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!