페르소나 프롬프팅이 실제로 효과적인 경우는 언제일까요? LLM에서 전문가 역할 주입의 검색 및 지표 분석
When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs
페르소나 프롬프팅은 대규모 언어 모델을 제어하는 데 널리 사용되지만, 그 실질적인 가치는 여전히 불분명합니다. 기존 연구에서는 종종 전체 점수를 사용하여 페르소나 프롬프팅을 평가하므로, 전문가 역할 주입이 응답 품질을 일관적으로 향상시키는지 아니면 단순히 다른 품질 측면에서 응답을 변화시키는지 판단하기 어렵습니다. 본 연구는 38개의 전문가 역할과 6가지 분야에 걸쳐 1,140개의 개방형 질문에 대한 4가지 프롬프팅 조건을 비교 분석하여 이 문제를 조사합니다: 역할 정보가 없는 기본적인 프롬프트, 일반적인 분야 전문가 프롬프트, 임베딩 기반의 역할 검색, 그리고 임베딩 검색과 LLM 기반 역할 선택을 결합한 하이브리드 검색 방법입니다. 전체 결과는 조건 간에 미미한 차이만 나타내는 것으로 보이지만, 지표 수준 분석에서는 전체 평균값이 가리는 일관적인 상호작용 관계가 드러납니다: 역할 프롬프팅은 전문성 깊이를 높이는 반면 명확성을 감소시킵니다. 이러한 효과는 보편적이라기보다는 조건에 따라 다르게 나타납니다. 역할 프롬프팅은 특히 의학 및 심리학 분야와 같이 구조화된 전문가 지식과 위험 정보 전달이 중요한 자문 관련 질문에서 가장 효과적입니다. 반면, 재무, 법률, 과학 및 기술 분야의 개념 설명 및 해설 관련 질문에서는 기본적인 프롬프트가 더 나은 성능을 보입니다. 또한 하이브리드 검색 방법이 임베딩 기반 역할 선택보다 훨씬 우수한 성능을 보이는 것으로 나타났습니다. 하지만 더 나은 역할 검색이라 할지라도 전문성 깊이와 명확성 간의 근본적인 상충 관계를 완전히 해소하지는 못합니다. 종합적으로, 본 연구 결과는 페르소나 프롬프팅이 전반적인 능력을 향상시키기보다는 응답 특성을 재구성하는 경향이 있으며, 그 효과를 이해하기 위해서는 다각적인 지표 평가가 필요하다는 것을 시사합니다.
Persona prompting is widely used to steer large language models, yet its practical value remains unclear. Prior work often evaluates persona prompting using aggregate scores, making it difficult to determine whether expert-role prompting consistently improves response quality or instead changes responses along different quality dimensions. We study this question through a controlled comparison of four prompting conditions across 1,140 open-ended questions spanning 38 expert roles and six domains: no role prompt, a generic domain-expert prompt, embedding-based role retrieval, and a hybrid retrieval method combining embedding search with LLM-based role selection. Aggregate results show only small overall differences between conditions. However, metric-level analysis reveals a consistent tradeoff that aggregate averages obscure: role prompting systematically increases expertise depth while reducing clarity. These effects are highly conditional rather than universal. Role prompting performs best on advisory questions and in domains such as medicine and psychology, where structured expert framing and risk communication are intrinsically valuable. In contrast, baseline prompting performs better on conceptual and explanatory questions in finance, legal, science, and technology domains, where concise plain-language explanation is more important. We further show that hybrid retrieval significantly improves over embedding-only role selection, although better role retrieval does not eliminate the broader expertise-depth versus clarity tradeoff. Overall, our findings suggest that persona prompting primarily reshapes response characteristics rather than broadly improving capability, and that multi-metric evaluation is necessary for understanding its effects.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.