개인화와 안전의 조화: 개인 맞춤형 LLM의 메커니즘, 위험 및 완화 방안
Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs
대규모 언어 모델(LLM)은 사용자의 선호도, 맥락 및 장기적인 기록에 맞춰 적응함으로써 점점 더 개인화된 상호작용을 가능하게 합니다. 그러나 개인화를 가능하게 하는 메커니즘은 기존 연구에서 체계적으로 다루지 않는 방식으로 안전 문제를 야기합니다. 기존의 리뷰는 일반적으로 개인화 또는 안전 중 하나에 초점을 맞추어, 이 두 가지의 교차 영역은 대부분 탐구되지 않았습니다. 본 연구에서는 개인 맞춤형 LLM에 대한 최초의 종합적이고 안전성을 고려한 리뷰를 제공합니다. 우리는 사용자 표현, 개인화 패러다임 및 평가라는 세 가지 측면으로 개인화를 분류하고, 통합된 안전 위험 분류 체계를 제시합니다. 표현 수준에서, 다양한 사용자 표현에서 발생하는 위험을 분석합니다. 주류 개인화 패러다임을 포괄적으로 검토하면서, 프롬프트, 검색 증강, 파라미터 미세 조정, 강화 학습, Mixture-of-Experts (MoE), 가지치기, 에이전트 프레임워크 및 다중 모드 개인화에 내재된 취약점을 명확히 하고, 모델 수명 주기 전반에 걸친 완화 전략을 종합합니다. 이러한 세분화된 위험 외에도, 개인화된 적응에서 발생하는 패러다임에 독립적인 안전 위험을 분석하고, 개인화된 데이터셋 및 평가 방법론을 요약합니다. OpenClaw의 사례 연구를 통해, 개인 맞춤형 에이전트 생태계에서의 배포 동향을 분석합니다. 우리의 분석 결과, 기존 연구에는 세 가지 구조적 결함이 존재한다는 것을 밝혀냈습니다: 안전성은 사용자 불변적인 방식으로 평가되며, 개인화 기술은 개별적으로 분석되지만 조합으로 고려되지 않으며, 평가 프레임워크는 예상치 못한 장기적인 위험을 포착할 수 없습니다. 본 연구에서는 개인 맞춤형 표현, 개인화 패러다임, 안전 위험, 방어 전략 및 평가 방법을 종합적으로 검토하여, 안전한 개인 맞춤형 LLM 개발을 위한 통합 프레임워크를 제시하고, 향후 연구의 주요 방향을 강조합니다.
Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. However, the mechanisms that enable personalization also expand the safety landscape in ways not systematically addressed by existing literature. Existing reviews typically focus either on personalization or safety, leaving their intersection largely unexplored. We present the first comprehensive, safety-aware review of personalized LLMs. We organize personalization along three dimensions-user representation, personalization paradigm, and evaluation-and introduce a unified taxonomy of safety risks. At the representation level, we analyze risks arising from diverse user representations. Across mainstream personalization paradigms, we delineate vulnerabilities inherent to prompting, retrieval augmentation, parameter fine-tuning, reinforcement learning, Mixture-of-Experts (MoE), pruning, agent frameworks, and multimodal personalization, and synthesize mitigation strategies across the model lifecycle. Beyond these fine-grained risks, we characterize paradigm-agnostic safety risks arising from personalized adaptation. We further summarize personalized datasets and evaluation methodologies. Through a case study of OpenClaw, we analyze deployment trends in personalized agent ecosystems. Our analysis reveals three structural inadequacies in existing research: safety is evaluated as user-invariant rather than relational, personalization techniques are analyzed in isolation rather than in composition, and evaluation frameworks cannot capture emergent long-term risks. By jointly examining personalized representations, personalization paradigms, safety risks, defenses, and evaluation methods, we provide a unified framework for developing safe personalized LLMs and highlight key directions for future research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.