로봇 학습 파이프라인을 손상시키기 위한 월드 모델 공격
Targeting World Models to Compromise Robot Learning Pipelines
최근 월드 모델은 로봇 학습 데이터 생성 또는 실제 환경 시뮬레이션을 위한 효율적인 도구로서 인기가 높아지고 있으며, 많은 연구에서 이를 로봇 학습 파이프라인에 통합하는 방안을 제시하고 있습니다. 본 논문에서는 이러한 월드 모델의 실용성에도 불구하고, 월드 모델이 로봇 학습 시스템 내에 은밀하고 효과적인 데이터 오염 경로를 제공하여, 겉보기에는 안전한 학습 데이터를 사용했음에도 불구하고 위험하거나 손상된 로봇 제어 정책이 배포될 수 있음을 보여줍니다. 기존의 데이터 오염 기법은 일반적으로 판매되거나 업로드된 데이터 세트에 직접적으로 위험한 경로를 삽입하는 반면, 본 연구에서는 눈에 띄는 안전한 원격 조작 데이터 세트에 악성 프롬프트 또는 유해한 상태 변화 동역학을 주입합니다. 이러한 공격은 월드 모델 입력으로 사용될 때만 활성화되며, 결과적으로 인공적인 위험한 로봇 학습 경로가 생성되고, 이는 최종적으로 안전하지 않거나 손상된 로봇 제어 정책으로 이어질 수 있습니다. 본 연구에서는 최첨단 액션 기반 및 텍스트 기반 월드 모델에 대한 공격의 효과를 입증했으며, 다운스트림 강화 학습(DRL) 정책에서 완전한 백도어를 구현하고 VLA 환경에서의 개념 증명을 보여주었습니다. 이러한 결과는 더욱 안전한 월드 모델 개발과 로봇 학습 시스템 내에서의 위치 재검토를 요구합니다.
World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. In contrast to traditional data poisoning techniques which directly implant dangerous trajectories into sold or uploaded datasets, our novel attack methods inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets which are only activated once fed through a world model as input. This can result in the generation of synthetic, dangerous robot training trajectories and subsequently unsafe or compromised robot policies. We demonstrate the effectiveness of our attacks against both state of the art action conditioned and text conditioned world models, showing a full end-to-end backdoor on a downstream DRL policy and a proof-of-concept for the VLA setting. Overall these findings necessitate research into more secure world models and reevaluating their position within the robot learning supply chain.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.