압박 하의 논리적 판단: 학습된 소프트 접두사를 활용한 삼단논법 안정성 진단
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
본 연구는 모델을 고정한 상태에서, 정확하게 레이블링된 삼단논법 추론 벤치마크에 '소프트 접두사'를 추가하여, 논리적 판단이 학습된 맥락에 어떻게 반응하는지 분석합니다. 소프트 접두사는 불투명한 연속 벡터로 구성되어 있으며, 우리는 이들이 논리 형태 및 인터페이스의 제어된 변화에 미치는 영향을 통해 그 특성을 규명합니다. 어떤 접두사가 성공적인지, 그리고 그 효과가 어떻게 일반화되는지를 연구함으로써, 학습된 맥락적 압력이 올바른 판단을 어떻게 무효화하고 모델의 논리적 안정성의 한계를 드러내는지 분석합니다. Qwen3.6-35B-A3B MoE, Qwen3-8B 및 Gemma 4 31B 모델에 대해, 학습된 접두사는 많은 올바른 답변을 변경시키며, 새로운 형태와 인터페이스 변화에도 효과적입니다. Qwen3.6 MoE 및 Gemma 모델에 대한 반복적인 테스트에서, 학습된 접두사는 16가지 모델-방향-분할 조합 모두에서 무작위 제어 그룹보다 37~99% 더 높은 성능을 보였습니다. Qwen3.6 MoE 모델의 오답률은 문장 구성 및 프롬프트 변경에 관계없이 72%에서 90% 사이로 유지되는 반면, Gemma 모델의 유효한 접두사는 무작위 접두사에 비해 54%에서 56%의 오답률을 보이는 데 그쳤습니다(무작위 접두사의 경우 1% 미만). 진단 테스트 결과, 주요 효과는 특정 답변에 대한 광범위한 선호도이며, 이는 고정된 심볼 강제 또는 작업 간에 안정적으로 전달되는 논리 연산 때문이 아닙니다. 이러한 편향의 형태는 모델마다 다릅니다. Qwen 모델에서는 간단한 점수 모델이 종종 어떤 판단이 뒤바뀔지 예측하지만, 그 변화 폭은 정확하게 예측하지 못하는 경우가 많습니다. 반면 Gemma의 전반적인 반응은 동일한 모델로 더 잘 근사됩니다. 이러한 결과는 학습된 소프트 접두사의 주요 효과가 특정 답변에 대한 광범위한 선호도라는 것을 보여주며, 나머지 반응은 모델별 논리적 안정성에 상당한 차이가 있음을 나타냅니다.
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model's logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, learned prefixes redirect many correct answers and remain effective across unseen forms and interface changes. In repeated tests with Qwen3.6 MoE and Gemma, they outperform paired random controls in all 16 model--direction--split comparisons by 37 to 99 percentage points. Qwen3.6 MoE flip rates remain between 72% and 90% across wording and prompt changes, while Gemma validity prefixes retain 54% to 56% flip compared with less than 1% for matched random prefixes. Diagnostic tests show that the dominant effect is a broad preference for one answer meaning rather than fixed-symbol forcing or a logical operation that transfers reliably between tasks. The form of this bias differs across models. In both Qwen models, simple score models often predict which judgments will flip but not how far their margins will move, whereas Gemma's overall response is more closely approximated by the same models. These results show that the dominant behavioral effect of successful soft prefixes is a broad answer preference, while the remaining response reveals substantial model-specific differences in logical stability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.