CtrlAttack: 확산 모델의 월드 모델 기반 제어에 대한 통합 공격
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
확산 기반 이미지-비디오(I2V) 모델은 시간적 동역학을 암묵적으로 포착하여 점차 월드 모델과 유사한 특성을 나타냅니다. 그러나 기존 연구는 주로 시각적 품질과 제어 가능성에 초점을 맞추었으며, 모델이 학습하는 상태 전이의 견고성은 상대적으로 연구가 부족했습니다. 이러한 간극을 메우기 위해, 본 연구는 I2V 모델의 취약점을 최초로 분석하고, 시간 제어 메커니즘이 새로운 공격 표면을 형성한다는 것을 밝혀냈으며, 다양한 공격 환경 하에서 이를 균일하게 모델링하는 데 따르는 어려움을 제시합니다. 이러한 분석을 바탕으로, 상태 진화 과정에 영향을 미치는 경로 제어 공격, 즉 CtrlAttack을 제안합니다. 구체적으로, 우리는 섭동을 저차원 속도장으로 표현하고, 시간 적분을 통해 연속적인 변위장을 구성하여 모델의 상태 전이에 영향을 미치면서 시간적 일관성을 유지합니다. 또한, 섭동을 관측 공간으로 매핑하여 본 연구 방법론이 화이트박스 및 블랙박스 공격 환경 모두에 적용될 수 있도록 합니다. 실험 결과, 낮은 차원의 강력한 정규화 제약 조건 하에서도, 본 연구 방법론은 공격 성공률(ASR)을 화이트박스 환경에서 90% 이상, 블랙박스 환경에서 80% 이상으로 증가시켜 시간적 일관성을 크게 파괴할 수 있다는 것을 보여줍니다. 동시에 FID 및 FVD 값의 변동을 각각 6과 130 이내로 유지하여, I2V 모델의 상태 동역학 수준에서 잠재적인 보안 위험을 드러냅니다.
Diffusion-based image-to-video (I2V) models increasingly exhibit world-model-like properties by implicitly capturing temporal dynamics. However, existing studies have mainly focused on visual quality and controllability, and the robustness of the state transition learned by the model remains understudied. To fill this gap, we are the first to analyze the vulnerability of I2V models, find that temporal control mechanisms constitute a new attack surface, and reveal the challenge of modeling them uniformly under different attack settings. Based on this, we propose a trajectory-control attack, called CtrlAttack, to interfere with state evolution during the generation process. Specifically, we represent the perturbation as a low-dimensional velocity field and construct a continuous displacement field via temporal integration, thereby affecting the model's state transitions while maintaining temporal consistency; meanwhile, we map the perturbation to the observation space, making the method applicable to both white-box and black-box attack settings. Experimental results show that even under low-dimensional and strongly regularized perturbation constraints, our method can still significantly disrupt temporal consistency by increasing the attack success rate (ASR) to over 90% in the white-box setting and over 80% in the black-box setting, while keeping the variation of the FID and FVD within 6 and 130, respectively, thus revealing the potential security risk of I2V models at the level of state dynamics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.