2606.20227v1 Jun 18, 2026 cs.AI

QMFOL: 양자화 가능한 단일 항 논리 테스트 케이스 생성을 통한 대규모 언어 모델 추론 성능 평가

QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation

Ling Shi
Ling Shi
Citations: 119
h-index: 5
Tianlong Yu
Tianlong Yu
Citations: 12
h-index: 2
Xinyi Zheng
Xinyi Zheng
Carnegie Mellon University
Citations: 428
h-index: 6
Yongxin Zhao
Yongxin Zhao
Citations: 35
h-index: 4
Lorenz F. Goette
Lorenz F. Goette
Citations: 3,406
h-index: 21
Kailong Wang
Kailong Wang
Citations: 2
h-index: 1

대규모 언어 모델(LLM)은 특히 연역적 추론 능력에서 상당한 발전을 이루었으며, 이는 중요한 의사 결정에 필수적입니다. 모델의 발전 속도에 맞춰 평가 벤치마크 또한 진화해야 합니다. 그러나 기존 벤치마크는 논리적 복잡성에 대한 세밀한 제어가 부족하고, 의미적 다양성과 논리적 일관성 사이의 균형을 맞추기 어렵다는 한계가 있습니다. 이러한 문제를 해결하기 위해, 우리는 QMFOL이라는 자동화된 프레임워크를 제안합니다. QMFOL은 양자화 가능하고 제어 가능한 복잡성을 가진 단일 항 논리 추론 과제를 생성하는 데 사용됩니다. 이 프레임워크는 논리적 구조를 구성할 때 결합 및 분리 패턴을 사용하여 추론의 깊이, 폭, 레이블 유형 및 오답 선택지를 정밀하게 제어할 수 있습니다. 이러한 구조는 LLM을 통해 자연어로 변환되며, 외부 증명기를 사용한 왕복 검증을 통해 논리적 일관성이 보장됩니다. 우리의 프레임워크를 기반으로, 우리는 다양한 논리적 및 의미적 차원을 포괄하는 2880개의 인스턴스와 960개의 구성으로 이루어진 QMFOLBench라는 벤치마크를 구축했습니다. 여섯 개의 대규모 추론 모델(LRM)과 두 개의 LLM에 대한 평가 결과, 논리적 복잡성이 증가함에 따라 성능이 저하되고 계산 비용이 증가한다는 것을 확인했습니다. 모델은 참(True) 레이블의 과제에서 거짓(False) 또는 알 수 없음(Unknown) 레이블의 과제보다 더 나은 성능을 보였으며, 의미 변형에도 민감하게 반응하는 것으로 나타났습니다. 전반적으로 QMFOL은 제어 가능한 복잡성을 가진 연역적 추론 벤치마크를 구축하기 위한 확장 가능하고 신뢰할 수 있는 접근 방식을 제공하며, 이를 통해 최신 언어 모델의 추론 능력을 보다 정확하게 평가할 수 있습니다.

Original Abstract

Large Language Models (LLMs) have made significant progress in reasoning, particularly in deductive reasoning, which is crucial for high-stakes decision-making. As models improve, evaluation benchmarks should evolve to keep pace. However, existing benchmarks lack fine-grained control over logical complexity and struggle to balance semantic diversity with logical consistency. To address these issues, we propose QMFOL, an automated framework for generating monadic first-order logic reasoning tasks with quantifiable and controllable complexity. It constructs formal logical structures using conjunction and disjunction patterns, enabling precise control over reasoning depth, width, label types, and distractors. These structures are then translated into natural language via LLMs, with logical consistency ensured through round-trip verification using an external prover. Based on our framework, we build QMFOLBench, a benchmark comprising 2880 instances with 960 configurations across diverse logical and semantic dimensions. Evaluations on six large reasoning models (LRMs) and two LLMs show that performance degrades and computational overhead increases with rising logical complexity. Models perform better on True-labeled tasks than on False or Unknown ones, and exhibit sensitivity to semantic variation. Overall, QMFOL offers a scalable and reliable approach for constructing deductive reasoning benchmarks with controllable complexity, enabling more precise evaluation of reasoning capabilities in modern language models.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!