FORGE: 심층 연구 에이전트에 대한 연구 경로 조작 공격
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
심층 연구 에이전트는 개방형 질문을 하위 작업으로 분해하고, 여러 단계에 걸쳐 웹에서 관련 정보를 검색하며, 긴 형식의 보고서를 작성합니다. 이러한 워크플로우는 계획 계층에 취약점을 야기하며, 검색 풀에 유입되는 악성 문서는 후속 질문의 방향을 조종하여 로컬 삽입이 보고서 전체에 영향을 미치도록 할 수 있습니다. 본 논문에서는 FORGE (Fabricated Orchestrated Reasoning chain for aGent Exploitation)라는 두 단계 공격을 제시합니다. 이는 문서 내부의 추론 위조와 문서 간 체인 연동을 결합하여 하위 작업 계획을 조작합니다. 또한, 감염된 보고서 주장을 인지적 유형에 따라 가중치를 부여하는 PRISM 지표를 소개하고, 재귀적인 후속 질문 생성을 초기 질문과 연결하는 경량 방어 기법인 Root Query Anchoring (RQA)을 제시합니다. 25개의 질문에 대해 Network FORGE는 5개의 삽입된 문서로 26.4%의 PRISM 값을 달성했으며, 재귀적인 합성 과정에서 악성 콘텐츠가 명백한 프레임에서 사실적 전제로 이동하는 현상 (depth migration)을 보였습니다. 10개의 질문으로 구성된 방어 실험에서는 RQA가 PRISM 값을 38.5%에서 18.3%로 감소시켰습니다.
Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This workflow creates a planning-layer poisoning surface: adversarial documents that enter the retrieval pool can steer follow-up questions and turn a local injection into report-level contamination. We present FORGE (Fabricated Orchestrated Reasoning chain for aGent Exploitation), a two-level attack that combines intra-document reasoning fabrication with inter-document chain coordination to hijack subtask planning. We further introduce the PRISM metric, which weights infected report claims by cognitive type, and Root Query Anchoring, a lightweight defense that ties recursive follow-up generation to the root query. Across 25 queries, Network FORGE reaches 26.4% PRISM with five injected documents and exhibits depth migration, in which recursive synthesis shifts poisoned content from overt framing into factual premises. On the 10-query defense subset, RQA (Root Query Anchoring) reduces PRISM from 38.5% to 18.3%.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.