2606.09125v1 Jun 08, 2026 cs.CR

멀티모달 대규모 언어 모델의 개인 정보 위험 분석: 작업별 취약점 및 완화 과제

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

Pingzhi Li
Pingzhi Li
Citations: 417
h-index: 9
Tiejin Chen
Tiejin Chen
Citations: 287
h-index: 8
Hua Wei
Hua Wei
Citations: 52
h-index: 3
Kaixiong Zhou
Kaixiong Zhou
Citations: 82
h-index: 5
Tianlong Chen
Tianlong Chen
Citations: 49
h-index: 3

텍스트 기반 대규모 언어 모델(LLM)에서 발생하는 개인 정보 위험은 잘 연구되어 왔으며, 특히 민감한 정보를 기억하고 유출하는 경향이 있습니다. 그러나 텍스트와 이미지를 함께 처리하는 멀티모달 대규모 언어 모델(MLLM)은 독특한 개인 정보 문제를 야기하며, 이는 아직 충분히 연구되지 않았습니다. 텍스트 기반 모델과 비교하여 MLLM은 이미지에 포함된 민감한 정보를 추출하고 노출할 수 있으며, 이는 새로운 개인 정보 위험을 초래합니다. 본 논문에서는 일부 MLLM이 개인 정보 침해에 취약하며, 이미지에 포함되거나 메모리에 저장된 민감한 데이터를 유출할 수 있음을 밝힙니다. 구체적으로, 본 연구에서는 (1) 다양한 멀티모달 작업 및 시나리오에서 개인 정보 위험을 평가하기 위해 설계된 포괄적인 데이터셋인 MM-Privacy를 소개하며, 여기서 노출 위험(Disclosure Risks)과 보존 위험(Retention Risks)을 정의합니다. (2) MM-Privacy를 사용하여 다양한 MLLM을 체계적으로 평가하고, 모델이 다양한 작업에서 민감한 데이터를 어떻게 유출하는지 보여줍니다. (3) 개인 정보 위험에 미치는 작업 불일치성의 역할을 분석하여 완화 전략의 시급성을 강조합니다. 본 연구 결과는 MLLM에서의 개인 정보 문제를 강조하며, 데이터 노출을 방지하기 위한 보안 조치의 필요성을 부각합니다. 본 연구에서 사용한 데이터셋 및 코드는 [링크]에서 확인할 수 있습니다.

Original Abstract

Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. However, Multi-modal Large Language Models (MLLMs), which process both text and images, introduce unique privacy challenges that remain underexplored. Compared to text-only models, MLLMs can extract and expose sensitive information embedded in images, posing new privacy risks. We reveal that some MLLMs are susceptible to privacy breaches, leaking sensitive data embedded in images or stored in memory. Specifically, in this paper, we (1) introduce MM-Privacy, a comprehensive dataset designed to assess privacy risks across various multi-modal tasks and scenarios, where we define Disclosure Risks and Retention Risks. (2) systematically evaluate different MLLMs using MM-Privacy and demonstrate how models leak sensitive data across various tasks, and (3) provide additional insights into the role of task inconsistency in privacy risks, emphasizing the urgent need for mitigation strategies. Our findings highlight privacy concerns in MLLMs, underscoring the necessity of safeguards to prevent data exposure. Our dataset and code can be found here.

16 Citations
0 Influential
4.5 Altmetric
38.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!