ImplicitBBQ: 특징 기반 신호를 활용한 대규모 언어 모델의 잠재적 편향성 평가
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
대규모 언어 모델은 인구학적 정보가 명시적으로 언급될 때 편향된 결과를 줄이는 경향이 있지만, 정보가 간접적으로 전달될 때 여전히 잠재적 편향성을 나타낼 수 있습니다. 기존의 벤치마크는 이름 기반의 지표를 사용하여 잠재적 편향성을 감지하는데, 이는 많은 사회적 인구 통계와 약한 연관성을 가지며, 나이, 사회경제적 지위와 같은 다른 차원으로 확장하기 어렵습니다. 본 연구에서는 특징 기반 신호를 활용하여 연령, 성별, 지역, 종교, 카스트, 사회경제적 지위 전반에 걸쳐 잠재적 편향성을 평가하는 질의응답 벤치마크인 ImplicitBBQ를 소개합니다. 11개의 모델을 평가한 결과, 모호한 맥락에서 나타나는 잠재적 편향성은 개방형 가중치 모델에서 나타나는 명시적 편향보다 6배 이상 높은 것으로 나타났습니다. 안전 프롬프트 및 체인 오브 소트 추론은 이러한 격차를 크게 줄이지 못하며, 잠재적 편향성을 84% 감소시키는 몇 가지 예시 프롬프트 사용에도 불구하고, 카스트 편향은 다른 모든 차원에 비해 4배 높은 수준을 유지합니다. 이러한 결과는 현재의 정렬 및 프롬프트 전략이 편향성 평가의 표면적인 측면만을 다루고 있으며, 문화적으로 기반된 고정관념적인 연관성을 해결하지 못한다는 것을 시사합니다. 본 연구에서는 모델 제공업체 및 연구자들이 잠재적인 완화 기술을 벤치마킹할 수 있도록 코드와 데이터셋을 공개합니다.
Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based proxies to detect implicit biases, which carry weak associations with many social demographics and cannot extend to dimensions like age or socioeconomic status. We introduce ImplicitBBQ, a QA benchmark that evaluates implicit bias through characteristic based cues, culturally associated attributes that signal implicitly, across age, gender, region, religion, caste, and socioeconomic status. Evaluating 11 models, we find that implicit bias in ambiguous contexts is over six times higher than explicit bias in open weight models. Safety prompting and chain-of-thought reasoning fail to substantially close this gap; even few-shot prompting, which reduces implicit bias by 84%, leaves caste bias at four times the level of any other dimension. These findings indicate that current alignment and prompting strategies address the surface of bias evaluation while leaving culturally grounded stereotypic associations largely unresolved. We publicly release our code and dataset for model providers and researchers to benchmark potential mitigation techniques.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.