계산적 선전(Computational Propaganda)을 통한 사전 학습 데이터 오염 가능성
Pretraining Data Can Be Poisoned through Computational Propaganda
사전 학습 데이터를 오염시키면 탐지하고 완화하기 어려운 유해한 동작을 언어 모델(LM)에 도입할 수 있습니다. 기존의 사전 학습 데이터 오염 연구는 주로 위키피디아와 같이 제한적인 특성을 가진 데이터 소스를 활용했으며, 오염된 데이터와 데이터 큐레이션 파이프라인 간의 상호 작용을 고려하지 않았습니다. 본 논문에서는 공개 토론 인터페이스를 이용한 기존의 웹 스케일 콘텐츠 주입 메커니즘을 통해, 위에서 언급한 제한적인 환경을 넘어 사전 학습 데이터에 대한 오염 공격이 가능하다는 것을 보여줍니다. 또한, 웹 크롤링 및 데이터 큐레이션 과정 이후 악성 콘텐츠가 포함되었는지 측정하기 위해, 웹 크롤링 기반 언어 모델 학습 데이터 내의 적대적 콘텐츠 포함 여사를 추정하는 새로운 분석 방법인 HalfLife를 소개합니다. 우리는 HalfLife를 사용하여 공개 토론 인터페이스를 통해 웹 스케일에서 사전 학습 코퍼스에 대한 오염 가능성을 탐구했습니다. 우리의 분석 결과는 사전 학습 데이터에 오염이 포함되었는지 추정하는 것의 중요성을 강조하며, 서드파티 웹페이지 콘텐츠가 언어 모델 사전 학습을 공격할 수 있는 잠재적인 경로임을 보여줍니다.
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.