에이전트 시스템의 지식 왜곡 방지: 비잔틴 내결함성을 갖춘 안전한 협업 RAG 프레임워크
Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
검색 증강 생성(RAG) 시스템은 대규모 언어 모델(LLM)의 환각 문제를 부분적으로 해결하지만, 동시에 지식 왜곡 공격에 대한 새로운 취약점을 야기합니다. 악의적인 사용자는 RAG 시스템이 제공하는 문서를 변조하여 LLM의 출력 결과를 조작할 수 있습니다. 이러한 위협에 대응하기 위해, 우리는 다중 소스 지식 검증 메커니즘을 활용하는 비잔틴 내결함성을 갖춘 협업 RAG 프레임워크인 SecureCollaRAG을 제안합니다. 우리 접근 방식은 에이전트 시스템이 동적 GNN 기반 신뢰도 점수를 통해 문서의 출처를 안전하게 검증하도록 하여, 은밀한 지식 왜곡 공격을 효과적으로 방지하고 중요한 도메인 지식의 무결성을 유지합니다. 광범위한 평가 및 형식적 분석을 통해 SecureCollaRAG이 비-IID 데이터 분포 하에서 공격자로부터 강력한 견고성을 유지함을 입증했습니다.
While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.