2607.28478v1 Jul 30, 2026 cs.CL

자동차 세차장에 걸어갈 의향이 있으신가요? - 거대 언어 모델의 상식 추론에서 나타나는 중요도 편향 분석

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

Zheng Wu
Zheng Wu
Citations: 99
h-index: 7
Zhuosheng Zhang
Zhuosheng Zhang
Citations: 434
h-index: 11
Chenhao Xue
Chenhao Xue
Citations: 4
h-index: 1
Shijie Zheng
Shijie Zheng
Citations: 0
h-index: 0
Yijie Lu
Yijie Lu
Citations: 5
h-index: 1
Cheng Yang
Cheng Yang
Citations: 0
h-index: 0

거대 언어 모델(LLM)이 복잡한 추론 과제에서 지속적으로 발전함에 따라, 입력된 명시적인 조건에 지나치게 우선순위를 부여하는 경향을 보입니다. 그러나 일상적인 상식 추론에서는 이러한 메커니즘이 심각한 취약점을 드러내는데, 이를 '중요도 편향(Salience Bias)'이라고 부릅니다. 모델은 쓸모없는 명시적 요소(예: 숫자 값)에 쉽게 영향을 받아, 특정 과제의 숨겨진 물리적 또는 상식적인 전제 조건을 무시하게 됩니다. 이 실패가 상식 지식 자체의 부족인지, 아니면 오해를 불러일으키는 과제 설정으로 인한 억압인지 여부는 중요한 질문입니다. 이를 조사하기 위해, 우리는 네 가지 유형의 함정을 포함하는 고품질 데이터셋인 'SaliTrap Benchmark'를 구축했습니다. 최첨단 LLM 12개를 평가한 결과, 모든 주류 모델이 상당한 수준의 중요도 편향을 보이는 것으로 나타났습니다. 편향의 심각성은 방해 요소의 밀도에 비례하며, 함정을 감지하는 것과 실제로 회피하는 것은 종종 별개로 나타났습니다. 중요한 점은, 과제 설정을 제거하고 동일한 모델을 다시 평가한 결과, 이는 지식 부족이 아닌 **지식 억압**으로 인한 실패라는 것을 보여주었습니다. 맥락 없는 지식 검증만으로도 위협적인 순응 실패의 90% 이상이 회복되었으며, 이는 필요한 상식이 본질적으로 존재하지만, 모델을 과도하게 순응적이고 불필요한 계산으로 유인하는 눈에 띄는 방해 요소에 의해 억제된다는 것을 보여줍니다. 이러한 진단을 바탕으로, 우리는 추가적인 학습 없이도 가벼운 추론 시간 프롬프팅만으로 상당한 성능 향상을 이룰 수 있음을 보여주었습니다. 우리의 연구 결과는 상식 추론 실패의 원인을 모델 자체의 능력 부족에서 벗어나, 정보 추출 방식에 있다는 것을 시사하며, SaliTrap을 이러한 문제점을 테스트할 수 있는 도구로 공개합니다. 관련 코드는 https://github.com/Wuzheng02/SaliTrap 에서 확인할 수 있습니다.

Original Abstract

As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90\% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at https://github.com/Wuzheng02/SaliTrap.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!