2606.09038v1 Jun 08, 2026 cs.AI

개인화와 안전의 조화: 개인 맞춤형 LLM의 메커니즘, 위험 및 완화 방안

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

Junlan Feng
Junlan Feng
Citations: 5
h-index: 1
Yanyan Luo
Yanyan Luo
Citations: 21
h-index: 2
Xue Han
Xue Han
Citations: 134
h-index: 4
Ruiqiao Bai
Ruiqiao Bai
Citations: 135
h-index: 5
Chunxu Zhao
Chunxu Zhao
Citations: 20
h-index: 2
Xinyi Huang
Xinyi Huang
Citations: 73
h-index: 4
Yitong Wang
Yitong Wang
Citations: 65
h-index: 3
Qian Hu
Qian Hu
Citations: 12
h-index: 2
Qing Wang
Qing Wang
Citations: 32
h-index: 3
Jie Liu
Jie Liu
Citations: 5
h-index: 2
Cong Geng
Cong Geng
Citations: 206
h-index: 4
Lehao Xing
Lehao Xing
Citations: 132
h-index: 3
Peng Hu
Peng Hu
Citations: 66
h-index: 3

대규모 언어 모델(LLM)은 사용자의 선호도, 맥락 및 장기적인 기록에 맞춰 적응함으로써 점점 더 개인화된 상호작용을 가능하게 합니다. 그러나 개인화를 가능하게 하는 메커니즘은 기존 연구에서 체계적으로 다루지 않는 방식으로 안전 문제를 야기합니다. 기존의 리뷰는 일반적으로 개인화 또는 안전 중 하나에 초점을 맞추어, 이 두 가지의 교차 영역은 대부분 탐구되지 않았습니다. 본 연구에서는 개인 맞춤형 LLM에 대한 최초의 종합적이고 안전성을 고려한 리뷰를 제공합니다. 우리는 사용자 표현, 개인화 패러다임 및 평가라는 세 가지 측면으로 개인화를 분류하고, 통합된 안전 위험 분류 체계를 제시합니다. 표현 수준에서, 다양한 사용자 표현에서 발생하는 위험을 분석합니다. 주류 개인화 패러다임을 포괄적으로 검토하면서, 프롬프트, 검색 증강, 파라미터 미세 조정, 강화 학습, Mixture-of-Experts (MoE), 가지치기, 에이전트 프레임워크 및 다중 모드 개인화에 내재된 취약점을 명확히 하고, 모델 수명 주기 전반에 걸친 완화 전략을 종합합니다. 이러한 세분화된 위험 외에도, 개인화된 적응에서 발생하는 패러다임에 독립적인 안전 위험을 분석하고, 개인화된 데이터셋 및 평가 방법론을 요약합니다. OpenClaw의 사례 연구를 통해, 개인 맞춤형 에이전트 생태계에서의 배포 동향을 분석합니다. 우리의 분석 결과, 기존 연구에는 세 가지 구조적 결함이 존재한다는 것을 밝혀냈습니다: 안전성은 사용자 불변적인 방식으로 평가되며, 개인화 기술은 개별적으로 분석되지만 조합으로 고려되지 않으며, 평가 프레임워크는 예상치 못한 장기적인 위험을 포착할 수 없습니다. 본 연구에서는 개인 맞춤형 표현, 개인화 패러다임, 안전 위험, 방어 전략 및 평가 방법을 종합적으로 검토하여, 안전한 개인 맞춤형 LLM 개발을 위한 통합 프레임워크를 제시하고, 향후 연구의 주요 방향을 강조합니다.

Original Abstract

Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. However, the mechanisms that enable personalization also expand the safety landscape in ways not systematically addressed by existing literature. Existing reviews typically focus either on personalization or safety, leaving their intersection largely unexplored. We present the first comprehensive, safety-aware review of personalized LLMs. We organize personalization along three dimensions-user representation, personalization paradigm, and evaluation-and introduce a unified taxonomy of safety risks. At the representation level, we analyze risks arising from diverse user representations. Across mainstream personalization paradigms, we delineate vulnerabilities inherent to prompting, retrieval augmentation, parameter fine-tuning, reinforcement learning, Mixture-of-Experts (MoE), pruning, agent frameworks, and multimodal personalization, and synthesize mitigation strategies across the model lifecycle. Beyond these fine-grained risks, we characterize paradigm-agnostic safety risks arising from personalized adaptation. We further summarize personalized datasets and evaluation methodologies. Through a case study of OpenClaw, we analyze deployment trends in personalized agent ecosystems. Our analysis reveals three structural inadequacies in existing research: safety is evaluated as user-invariant rather than relational, personalization techniques are analyzed in isolation rather than in composition, and evaluation frameworks cannot capture emergent long-term risks. By jointly examining personalized representations, personalization paradigms, safety risks, defenses, and evaluation methods, we provide a unified framework for developing safe personalized LLMs and highlight key directions for future research.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!