2602.02462v1 Feb 02, 2026 cs.CL

대규모 언어 모델에서 콘텐츠 불변 추론을 위한 추상화된 활성화 공간

Abstract Activation Spaces for Content-Invariant Reasoning in Large Language Models

M. Valentino
M. Valentino
Citations: 22
h-index: 3
Gabriele Maraia
Gabriele Maraia
Citations: 3
h-index: 1
F. Zanzotto
F. Zanzotto
Citations: 694
h-index: 15
Leonardo Ranaldi
Leonardo Ranaldi
Citations: 736
h-index: 16

대규모 언어 모델(LLM)은 종종 삼단논법 추론에서 연역적 판단을 수행하는 데 어려움을 겪으며, 의미적 타당성과 형식적 타당성을 체계적으로 혼동하는 현상, 즉 '콘텐츠 효과'를 나타냅니다. 이러한 편향은 모델이 단계별 설명을 생성하더라도 지속되며, 이는 중간 단계의 추론 과정이 답변에 영향을 미치는 동일한 의미적 단축키를 포함할 수 있음을 시사합니다. 최근의 접근 방식은 추론 시간 동안의 구조적 제약을 강화하여 이 문제를 완화하려고 시도하며, 이는 추상적인 중간 표현을 장려하거나 모델의 내부 계산에 직접 개입하는 방식을 포함합니다. 그러나 의미적 간섭을 안정적으로 억제하는 것은 여전히 해결해야 할 과제입니다. 우리는 형식적 추론이 의미적 내용에 덜 민감하도록 만들기 위해, 구조적 추론과 어휘적 의미를 명시적으로 분리하는 추상화 기반 추론 프레임워크를 소개합니다. 우리는 콘텐츠가 풍부한 삼단논법과 추상적인 삼단논법을 쌍으로 구성하고, 모델이 추상적인 입력에 대해 생성하는 활성화를 사용하여 추상적 추론 공간을 정의합니다. 그런 다음, 콘텐츠에 따라 조건화된 잔류 스트림 상태에서 이 공간과 일치하는 표현을 예측하는 가벼운 추상화 모델을 학습하고, 이러한 예측을 순방향 패스 과정에서 다층 개입을 통해 통합합니다. 교차 언어 전이(cross-lingual transfer)를 테스트 환경으로 사용하여, 추상화에 따른 조향(steering)이 콘텐츠 기반 오류를 줄이고 타당성 민감한 성능을 향상시킴을 보여줍니다. 우리의 연구 결과는 활성화 수준의 추상화를 대규모 언어 모델에서 형식적 추론의 견고성을 향상시키는 확장 가능한 메커니즘으로 제시합니다.

Original Abstract

Large Language Models (LLMs) often struggle with deductive judgment in syllogistic reasoning, systematically conflating semantic plausibility with formal validity a phenomenon known as content effect. This bias persists even when models generate step-wise explanations, indicating that intermediate rationales may inherit the same semantic shortcuts that affect answers. Recent approaches propose mitigating this issue by increasing inference-time structural constraints, either by encouraging abstract intermediate representations or by intervening directly in the model's internal computations; however, reliably suppressing semantic interference remains an open challenge. To make formal deduction less sensitive to semantic content, we introduce a framework for abstraction-guided reasoning that explicitly separates structural inference from lexical semantics. We construct paired content-laden and abstract syllogisms and use the model's activations on abstract inputs to define an abstract reasoning space. We then learn lightweight Abstractors that, from content-conditioned residual-stream states, predict representations aligned with this space and integrate these predictions via multi-layer interventions during the forward pass. Using cross-lingual transfer as a test bed, we show that abstraction-aligned steering reduces content-driven errors and improves validity-sensitive performance. Our results position activation-level abstraction as a scalable mechanism for enhancing the robustness of formal reasoning in LLMs against semantic interference.

5 Citations
0 Influential
8 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!