2602.10117v3 Feb 10, 2026 cs.LG

사각지대의 편향: LLM이 언급하지 못하는 내용을 탐지하기

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

Iv'an Arcuschin
Iv'an Arcuschin
Citations: 282
h-index: 6
David Chanin
David Chanin
Citations: 341
h-index: 6
Adrià Garriga-Alonso
Adrià Garriga-Alonso
Citations: 4,097
h-index: 11
Oana-Maria Camburu
Oana-Maria Camburu
Citations: 330
h-index: 10

대규모 언어 모델(LLM)은 종종 그럴듯해 보이는 추론 과정을 제공하지만, 내부적인 편향을 숨길 수 있습니다. 우리는 이러한 편향을 *명시되지 않은 편향(unverbalized biases)*이라고 부릅니다. 따라서 모델의 명시된 추론 과정을 통해 모델을 모니터링하는 것은 신뢰할 수 없으며, 기존의 편향 평가 방법은 일반적으로 미리 정의된 범주와 수작업으로 만들어진 데이터 세트를 필요로 합니다. 본 연구에서는 작업별 명시되지 않은 편향을 탐지하기 위한 완전 자동화된 블랙박스 파이프라인을 소개합니다. 주어진 작업 데이터 세트를 기반으로, 이 파이프라인은 LLM 자동 평가기를 사용하여 잠재적인 편향 개념을 생성합니다. 그런 다음, 파이프라인은 각 개념을 점진적으로 더 큰 입력 샘플에 대해 테스트하여 긍정적 및 부정적 변형을 생성하고, 다중 검정 및 조기 종료를 위한 통계적 기법을 적용합니다. 모델의 추론 과정에서 언급되지 않으면서 통계적으로 유의미한 성능 차이를 보이는 개념은 명시되지 않은 편향으로 표시됩니다. 우리는 제안하는 파이프라인을 세 가지 의사 결정 작업(채용, 대출 승인, 대학 입학)에 대해 7개의 LLM을 사용하여 평가했습니다. 우리의 기술은 이러한 모델에서 이전에 알려지지 않았던 편향을 자동으로 발견합니다(예: 스페인어 구사 능력, 영어 능력, 글쓰기 형식). 동시에, 이 파이프라인은 기존 연구에서 수동으로 식별된 편향(성별, 인종, 종교, 민족)을 검증합니다. 더 넓은 관점에서, 제안하는 접근 방식은 작업별 편향을 자동으로 발견하는 실용적이고 확장 가능한 방법을 제공합니다.

Original Abstract

Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We call these *unverbalized biases*. Monitoring models via their stated reasoning is therefore unreliable, and existing bias evaluations typically require predefined categories and hand-crafted datasets. In this work, we introduce a fully automated, black-box pipeline for detecting task-specific unverbalized biases. Given a task dataset, the pipeline uses LLM autoraters to generate candidate bias concepts. It then tests each concept on progressively larger input samples by generating positive and negative variations, and applies statistical techniques for multiple testing and early stopping. A concept is flagged as an unverbalized bias if it yields statistically significant performance differences while not being cited as justification in the model's CoTs. We evaluate our pipeline across seven LLMs on three decision tasks (hiring, loan approval, and university admissions). Our technique automatically discovers previously unknown biases in these models (e.g., Spanish fluency, English proficiency, writing formality). In the same run, the pipeline also validates biases that were manually identified by prior work (gender, race, religion, ethnicity). More broadly, our proposed approach provides a practical, scalable path to automatic task-specific bias discovery.

4 Citations
0 Influential
5.5 Altmetric
31.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!