2607.18622v1 Jul 21, 2026 cs.CR

CPInj: 텍스트 기반 협력적 프롬프트 최적화 시스템의 프롬프트 주입 위험 분석

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

D. Pandya
D. Pandya
Citations: 247
h-index: 10
Xinting Liao
Xinting Liao
Citations: 1
h-index: 1
Xiaoxiao Li
Xiaoxiao Li
Citations: 1
h-index: 1
Behnoosh Zamanlooy
Behnoosh Zamanlooy
Citations: 98
h-index: 5
Masoumeh Shafieinejad
Masoumeh Shafieinejad
Citations: 245
h-index: 7
D. B. Emerson
D. B. Emerson
Citations: 4
h-index: 2
Ruinan Jin
Ruinan Jin
Citations: 105
h-index: 5

텍스트 기반 협력적 프롬프트 최적화(TCPO)는 Textgrad (Yuksekgonul et al., 2025)를 분산 환경으로 확장하여, 여러 클라이언트가 대규모 언어 모델(LLM)에 대한 프롬프트를 공동으로 개선하면서 각자의 데이터를 로컬에서 유지하도록 합니다. TCPO는 자유 형식의 텍스트 업데이트 및 집계를 사용하는데, 이는 새로운 공격 취약점을 야기합니다. 즉, 악성 명령이 로컬 프롬프트에 주입되어 서버 측 프롬프트 집계 과정을 통해 전파될 수 있습니다. 기존의 프롬프트 주입 공격과는 달리, TCPO를 공격하는 것은 TCPO의 협력적 최적화 루프를 목표로 합니다. 이 환경은 더 어렵습니다. 왜냐하면 악성 명령은 집계를 통과하고, 이후의 정상적인 프롬프트 최적화를 견뎌내며, 서버 측 방어 시스템을 회피해야 하기 때문입니다. 이러한 위험을 드러내기 위해, 우리는 CPInj라는 협력적 프롬프트 주입 공격을 제안합니다. CPInj는 악성 명령으로 구성된 오염된 글로벌 프롬프트를 생성하고, 하위 작업의 성능을 저하시키며, 정상 클라이언트에 의한 프롬프트 최적화를 통해 제거되지 않도록 설계되었으며, 서버 측의 고급 탐지 기반 방어를 회피합니다. 우리는 현재의 방어 방법이 CPInj에 대해 효과가 없음을 확인했습니다. 이 공격을 완화하기 위해, 우리는 악성 명령을 정제하고 TCPO의 유용성을 부분적으로 복구하는 방어 중심 집계 방법인 APAgg를 추가로 제안합니다. 우리는 세 가지 LLM 패밀리와 수학, 논리 및 의학 분야의 다섯 가지 추론 작업을 포함하여 광범위한 실험을 수행했습니다. 결과는 우리의 제안된 공격이 TCPO에서 중요한 취약점을 드러낸다는 것을 보여줍니다. 우리는 완화에 대한 첫걸음을 내딛었지만, 이 공격은 여전히 매우 효과적이며 완전히 해결되지 않았으므로, TCPO를 위한 더욱 강력한 방어가 필요합니다.

Original Abstract

Textual Collaborative Prompt Optimization (TCPO) extends Textgrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multiple clients to jointly improve prompts for large language models (LLMs) while keeping their data locally. Its reliance on free-form textual updating and aggregation introduces a new and largely unexplored attack surface, i.e., malicious instructions can be injected into local prompts and propagated through server-side prompt aggregation. Unlike conventional prompt injection attacks, attacking TCPO targets the collaborative optimization loop in TCPO. This setting is more challenging because malicious instructions must survive aggregation, persist through subsequent benign prompt optimization, and evade server-side defenses. To expose this risk, we propose CPInj, a collaborative prompt injection attack that contaminates the aggregated global prompt with malicious instructions, degrades downstream task performance, resists purification by prompt optimization on benign clients, and evades advanced detection-based defenses on the server. We find that current defense methods are ineffective against CPInj. To mitigate this attack, we further propose a defense-oriented aggregation method, i.e., APAgg, which purifies malicious instructions and partially recovers TCPO utility. We conduct extensive experiments across three LLM families and five reasoning tasks in math, logic, and medicine. The results demonstrate that our proposed attack reveals a critical vulnerability in TCPO. Although we take a first step toward mitigation, the attack remains highly effective and far from fully resolved, calling for more robust defense for TCPO.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!