MuPPET: 다자간 대화에서 LLM 어시스턴트의 맥락적 개인정보 보호를 위한 벤치마크
MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations
LLM 에이전트는 점점 더 많은 다자 환경에 배포되어 개별 사용자를 대신하여 민감한 개인 정보를 처리하며, 예를 들어 그룹 채팅에서 그러한 역할을 수행합니다. 이러한 에이전트가 사적인 정보를 공개할 경우, 해당 정보는 한 번에 모든 그룹 구성원에게 전달됩니다. 이러한 위험은 1:1 환경보다 제어하기 어렵습니다. 왜냐하면 각 사적인 정보 조각이 그룹의 모든 수신자에게 적절해야 하기 때문입니다. 그러나 기존의 맥락적 개인정보 보호 벤치마크는 단일 대화 상대 환경만을 고려하여, 다자간 환경에서의 개인정보 침해 위험을 측정하지 못합니다. 우리는 다자간 대화에서 맥락적 개인정보 보호를 위한 벤치마크인 MuPPET (Multi-Party Privacy Exposure Testing)을 소개합니다. 우리의 실험 결과는 모델이 다자간 환경에서 1:1 평가보다 훨씬 더 많은 정보를 유출한다는 것을 보여줍니다. 최첨단 모델은 취약하며, 특히 민감한 데이터를 로컬 환경에 배포하는 데 자주 사용되는 소규모 오픈 웨이트 모델의 취약성이 더욱 심각합니다. 기존의 맥락적 개인정보 보호 기술은 부분적인 보호만 제공하고, 성능을 저하시키며, 근본적인 사용자 추적 문제를 해결하지 못합니다.
LLM agents are increasingly deployed in multi-party environments, handling sensitive personal data on behalf of individual users, for instance in group chats. When such an agent discloses private information, it reaches every group member at once. This risk is structurally harder to control than in one-to-one settings, as every piece of private information must be appropriate for every recipient in the group. Yet all existing contextual privacy benchmarks consider only single-interlocutor settings, leaving multi-party privacy risks unmeasured. We introduce MuPPET (Multi-Party Privacy Exposure Testing), a benchmark for contextual privacy in multi-party conversations. Our experiments show that models leak substantially more in multi-party settings than one-to-one evaluations suggest. Frontier models are vulnerable, and smaller open-weights models, often preferred for local deployment with sensitive data, even more so. Existing contextual privacy defences offer only partial protection, degrade utility, and do not resolve the underlying party-tracking problem.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.