2608.06955v1 Aug 07, 2026 cs.AI

대규모 언어 모델의 비평적 평가 선호도: 영화 선호도 추출 연구를 통한 증거

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

Aaron Shaw
Aaron Shaw
Citations: 10
h-index: 1
Jonghyun Jee
Jonghyun Jee
Citations: 1
h-index: 1

대규모 언어 모델(LLM)은 영화, 책, 음악 등 다양한 분야에 대한 인간의 평가 표현을 포함하는 데이터셋으로 학습됩니다. 하지만 LLM이 체계적으로 평가 위계를 재현하는지 여부는 아직 명확하지 않습니다. LLM 내 문화적 편향성에 대한 기존 연구는 상반된 기대를 제시합니다. 모델은 인터넷 텍스트에서 나타나는 인기 신호를 반영할 수도 있지만, 비평 담론에 내재된 권위를 재현할 수도 있습니다. 본 연구에서는 앤트로픽(Anthropic), OpenAI, 알리바바(Alibaba), 미스트랄(Mistral)의 네 가지 계열에서 개발된 8개의 모델을 사용하여 200편의 영화를 대상으로 한 평가 연구를 통해 이 질문에 대한 답을 찾고자 합니다. 벤치마크 데이터셋은 비평적으로 높은 평가를 받았지만 상업적 성공이 낮은 영화, 상업적으로 성공했지만 비평적으로 인정받지 못한 영화, 그리고 비평적 명성과 상업적 성공 모두를 가진 영화로 나뉩니다. 각 모델당 약 20,000개의 쌍대 비교 데이터를 브래들리-테리(Bradley--Terry) 추정법을 사용하여 분석한 결과, 모든 모델에서 일관된 비평적 평가 선호도 경향이 나타났습니다. 즉, 상업적으로 성공하지 못했지만 비평적으로 높은 평가를 받은 영화가 상업적으로 성공했지만 비평적으로 인정받지 못한 영화보다 더 자주 선택되었습니다. 이러한 패턴은 각 계열 내에서 모델 규모가 커짐에 따라 더욱 뚜렷해집니다. 또한, 중첩된 OLS 회귀 분석 결과, 평가적 성향, 대중적 인지도, 그리고 인기 정도가 선호도 결정에 명확하게 기여하는 것으로 나타났습니다. 대중적 인지도를 고려하면 모델이 비평적 명성만 있는 영화보다 상업적 성공과 비평적 명성을 모두 가진 영화를 더 선호하는 경향이 반전됩니다. 또한, 인기 정도를 추가적으로 고려하면 상업적 성공만 있는 영화의 불리함이 상당히 완화되는 것으로 나타났습니다. 마지막으로, 평가 지향적 및 추천 지향적인 프롬프트 형식이 서로 다른 순위를 생성한다는 점은, 비평적 평가 선호도가 실제 LLM 운영 환경에서 간접적으로 나타날 수 있음을 시사합니다.

Original Abstract

Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over commercially successful yet critically unrecognized ones. This pattern grows with model scale within each family. In addition, nested OLS regression analyses show that evaluative orientation, public visibility, and popular reception distinctly help explain preferences. Adjusting for public visibility reverses the models' preference for dual-legitimacy films over critical acclaim-only films, while additionally accounting for popular reception attenuates much of the disadvantage of films with commercial success only. Finally, evaluative and recommendation-oriented prompt framings produce divergent rankings, suggesting that critical acclaim orientation may manifest indirectly in real-world LLM deployments.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!