동일한 공식, 다른 의미: 언어 모델이 양상 논리 명세를 따르는가?
Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?
필연성과 가능성에 대한 추론은 세계 간의 접근성에 대한 가정과 각 세계에 존재하는 객체에 대한 가정에 의존합니다. 따라서 동일한 추론이라도 하나의 양상 체계에서는 성립할 수 있지만 다른 체계에서는 그렇지 않을 수 있습니다. 이러한 문제에 대해 언어 모델을 평가하려면, 모델의 판단이 명시된 의미론을 따르는지, 단순히 익숙한 논리를 따르는지를 확인해야 합니다. 우리는 동일한 전제와 결론을 갖지만 프레임 또는 도메인 조건이 다른 쌍으로 구성된 양상 문제를 만들었습니다. 자동 추론은 반대 결과를 보입니다. 균형 잡힌 문제 세트는 의미론적 조건만으로는 답을 얻을 수 없도록 설계되었습니다. 이러한 문제에서 최근 개발된 모델 5개 중 4개가 직접적인 프롬프트 사용 시, 조건 정보만을 사용하는 기준점보다 낮은 성능을 보였습니다. 그러나 추론 모드를 활성화하면 DeepSeek V4 Flash 모델의 정확도가 동일한 프롬프트를 사용하여 4.4%에서 88.1%로 크게 향상되었습니다. 따라서 명시된 양상 의미론을 따르는 것은 추론 모드뿐만 아니라 모델 자체에 따라 크게 달라집니다. 프레임 조건이 생략되면, 모델들은 종종 일관성을 보이지만, 각 모델은 서로 다른 익숙한 논리에 가장 잘 부합합니다. 우리는 사용된 공식, 정답 데이터, 반례 및 모델 응답을 공개합니다.
Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same inference may therefore hold under one modal system and fail under another. Evaluating language models on such problems requires testing whether their judgments follow the stated semantics rather than a familiar logic. We construct paired modal problems with identical premises and conjecture but different frame or domain conditions; automated reasoning verifies opposite labels. A balanced core prevents the semantic condition alone from revealing the answer. On this core, four of five recent models perform below the condition-only baseline under direct prompting. Yet enabling reasoning mode raises DeepSeek V4 Flash from 4.4% to 88.1% on unchanged prompts. Following stipulated modal semantics thus depends strongly on inference mode as well as model identity. When frame conditions are omitted, models often agree but fit different familiar logics best. We release the formulas, oracle artifacts, countermodels, and responses.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.