2608.06977v1 Aug 07, 2026 cs.CL

우리 편향을 확인하는가? 대규모 언어 모델의 기능, 위험 및 사회적 영향 평가

Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models

Polina Tsvilodub
Polina Tsvilodub
Citations: 61
h-index: 5
Mudar Adas
Mudar Adas
Citations: 0
h-index: 0
Michael Franke
Michael Franke
Citations: 44
h-index: 4
Martin V. Butz
Martin V. Butz
Citations: 24
h-index: 3

대규모 언어 모델(LLM)은 프롬프트 구성에 민감하며, 이는 학습 데이터 또는 이전 프롬프트에서 비롯된 패턴을 반영합니다. 본 연구에서는 LLM이 프롬프트에 표현된 사용자의 편향을 얼마나 강화하는지 조사하고, 암묵적인 프레임 효과와 명시적인 프롬프트 조작 사이의 경계를 검토합니다. 구체적으로, 우리는 모델이 특정 입장을 지지하거나 반박하도록 유도하는 직접적이고 간접적인 프롬프트에 LLM이 얼마나 취약한지를 평가합니다. 본 연구에서는 10개의 주제를 아우르는 의견 기반 및 사실 영역에서 총 160개의 다양한 프롬프트를 사용하여 6개의 LLM을 평가했습니다. 프롬프트는 프롬프트 전략, 지지 또는 반박 지시사항, 프롬프트 극성, 사용자의 표현된 신념 및 주제 영역에 대해 체계적으로 변형되었습니다. 연구 결과는 LLM이 사실적인 맥락에서도 프롬프트 구성에 따라 응답을 체계적으로 조정한다는 것을 보여줍니다. 이는 프롬프트 구성이 모델 응답의 사실적 일관성을 능가할 수 있음을 시사합니다. 전반적으로, 본 연구의 결과는 LLM의 조작 가능성의 범위와 경계를 명확히 합니다. 또한, 본 연구 결과는 LLM이 미묘한 사용자 편향을 강화하고, 사실적으로 안정적인 응답이 필요하지만 명시적인 프롬프트 조작에 취약할 수 있음을 시사합니다.

Original Abstract

It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases expressed in the prompts and examine the boundary between implicit framing effects and explicit prompt manipulation. Specifically, we evaluate how susceptible LLMs are to direct and suggestive prompts that encourage models to support or challenge particular positions. We evaluate six LLMs using 160 distinct prompts spanning ten topics across opinion-based and factual domains. The prompts systematically vary in prompting strategy, support versus challenge instructions, prompt polarity, users' expressed beliefs, and topic domain, spanning both opinion-based and factual questions. Our results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts. This suggests that prompt framing can outweigh factual consistency in model responses. Overall, our findings delineate the extent and boundaries of LLM manipulability. Furthermore, the results imply that LLMs can reinforce subtle user biases and are susceptible to explicit prompt manipulation even in domains where responses should remain factually stable.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!