2602.04398v1 Feb 04, 2026 cs.CL

양방향 편향 설명: 프롬프트를 수정하지 않고 대규모 언어 모델의 편향 제거

Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying Prompts

Yujie Lin
Yujie Lin
Citations: 147
h-index: 7
Kunquan Li
Kunquan Li
Citations: 201
h-index: 3
Yixuan Liao
Yixuan Liao
Citations: 110
h-index: 5
Jinsong Su
Jinsong Su
Citations: 23
h-index: 4
Xiaoxin Chen
Xiaoxin Chen
Citations: 1,803
h-index: 15

대규모 언어 모델(LLM)은 다양한 자연어 처리 작업에서 뛰어난 성능을 보여주었습니다. 그러나 이들의 출력은 종종 사회적 편향을 나타내어 공정성에 대한 우려를 야기합니다. 기존의 편향 제거 방법, 예를 들어 추가 데이터 세트에 대한 미세 조정 또는 프롬프트 엔지니어링은 확장성 문제에 직면하거나 다중 턴 상호 작용에서 사용자 경험을 저해할 수 있습니다. 이러한 문제점을 해결하기 위해, 우리는 미세 조정이나 프롬프트 수정 없이 대규모 언어 모델에서 스테레오타입을 유발하는 단어를 감지하고 신경망 수준의 편향을 설명하는 프레임워크를 제안합니다. 우리의 프레임워크는 먼저 인구 집단 간의 비교 분석을 통해 스테레오타입을 유발하는 형용사와 명사를 식별합니다. 그런 다음, 통합 그래디언트를 기반으로 한 두 가지 설명 전략을 사용하여 편향된 행동을 특정 뉴런에 연결합니다. 마지막으로, 투영 레이어에서 해당 뉴런의 활성값을 직접 수정하여 편향을 완화합니다. 널리 사용되는 세 가지 대규모 언어 모델에 대한 실험 결과, 우리의 방법은 전체 모델 성능을 유지하면서 효과적으로 편향을 줄이는 것을 보여줍니다. 코드는 다음 github 링크에서 확인할 수 있습니다: https://github.com/XMUDeepLIT/Bi-directional-Bias-Attribution.

Original Abstract

Large language models (LLMs) have demonstrated impressive capabilities across a wide range of natural language processing tasks. However, their outputs often exhibit social biases, raising fairness concerns. Existing debiasing methods, such as fine-tuning on additional datasets or prompt engineering, face scalability issues or compromise user experience in multi-turn interactions. To address these challenges, we propose a framework for detecting stereotype-inducing words and attributing neuron-level bias in LLMs, without the need for fine-tuning or prompt modification. Our framework first identifies stereotype-inducing adjectives and nouns via comparative analysis across demographic groups. We then attribute biased behavior to specific neurons using two attribution strategies based on integrated gradients. Finally, we mitigate bias by directly intervening on their activations at the projection layer. Experiments on three widely used LLMs demonstrate that our method effectively reduces bias while preserving overall model performance. Code is available at the github link: https://github.com/XMUDeepLIT/Bi-directional-Bias-Attribution.

5 Citations
0 Influential
32.993061443341 Altmetric
17.9 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!