대규모 언어 모델을 이용한 성별 추론에서 발생하는 맥락 불변성 실패
Failure of contextual invariance in gender inference with large language models
기존의 평가 방식은 대규모 언어 모델(LLM)의 결과가 작업의 맥락적으로 동일한 표현 하에서 안정적이라고 가정합니다. 본 연구에서는 성별 추론이라는 환경에서 이러한 가정을 검증합니다. 통제된 대명사 선택 작업을 통해, 우리는 최소한의 이론적으로 무의미한 담론 맥락을 도입하고, 이로 인해 모델 결과에 큰 변화가 발생하는 것을 확인했습니다. 비맥락화된 환경에서 나타나는 문화적 성별 고정관념과의 상관관계는 맥락이 도입되면 약화되거나 사라지는 반면, 이론적으로 관련 없는 특징(예: 무관련 대명사의 성별)이 모델의 행동을 예측하는 가장 중요한 요소가 됩니다. '기본적으로 맥락 의존적'이라는 분석 결과, 모델에 따라 19~52%의 경우, 개별 결과에 대한 모든 주변 효과를 고려한 후에도 이러한 맥락 의존성이 지속되며, 이는 단순한 대명사 반복 때문이 아님을 보여줍니다. 이러한 결과는 LLM의 결과가 거의 동일한 구문 구조에서도 맥락 불변성을 위반한다는 것을 보여주며, 이는 편향성 평가 및 고위험 환경에서의 활용에 중요한 함의를 가집니다.
Standard evaluation practices assume that large language model (LLM) outputs are stable under contextually equivalent formulations of a task. Here, we test this assumption in the setting of gender inference. Using a controlled pronoun selection task, we introduce minimal, theoretically uninformative discourse context and find that this induces large, systematic shifts in model outputs. Correlations with cultural gender stereotypes, present in decontextualized settings, weaken or disappear once context is introduced, while theoretically irrelevant features, such as the gender of a pronoun for an unrelated referent, become the most informative predictors of model behaviour. A Contextuality-by-Default analysis reveals that, in 19--52\% of cases across models, this dependence persists after accounting for all marginal effects of context on individual outputs and cannot be attributed to simple pronoun repetition. These findings show that LLM outputs violate contextual invariance even under near-identical syntactic formulations, with implications for bias benchmarking and deployment in high-stakes settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.