2510.27544v3 Oct 31, 2025 cs.AI

TempoBench: 인과적 추론은 원인 규명 없이 실행되는 것일 뿐, 단지 시뮬레이션이다

TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation

Nikolaus Holzer
Nikolaus Holzer
Citations: 27
h-index: 3
William Fishell
William Fishell
Citations: 1
h-index: 1
Baishakhi Ray
Baishakhi Ray
Citations: 122
h-index: 6
Mark Santolucito
Mark Santolucito
Citations: 375
h-index: 11

장기적인 추론 과정을 최적화하도록 설계된 현재의 학습 패러다임 덕분에 대규모 언어 모델(LLM)은 패턴 매칭 및 추론 과정의 순방향 시뮬레이션에서 뛰어난 성능을 보이지만, 반사실적 인과 관계 이해 및 추론에서는 상대적으로 미흡합니다. 본 논문에서는 실행 경로 전반에 걸친 반사실적 인과적 귀인 과정을 명확하게 분리하는 최초의 형식적으로 검증 가능한 시간 기반 벤치마크인 TempoBench를 소개합니다. TempoBench를 통해 LLM이 인과 관계 추론 문제 해결 시 근본적으로 무차별적인 시뮬레이션 기반 추론에 의존한다는 것을 보여줍니다. 합성된 결정적 Mealy 머신으로 구축된 TempoBench는 제어 가능한 복잡성과 증명 가능한 정확한 인과 관계 레이블을 갖춘 무한히 확장 가능한 경로 기반 인과 관계 추론 문제 모음을 제공합니다. 최첨단 모델은 시스템의 순방향 시뮬레이션에서 96%의 단계별 정확도를 달성하지만, 관찰된 출력에 필요한 입력이 무엇인지 질문하면 정확도가 32%로 떨어지는데, 이를 우리는 SIM/MIN 격차라고 부릅니다. 우리의 연구 결과는 LLM이 최소한으로 필요한 원인을 안정적으로 식별할 수 없으며, 종종 '가능한 입력'과 '필요한 원인'을 혼동하여 어떤 입력이 필요하지 않았는지 이해하는 능력이 부족함을 보여줍니다. 이러한 실패는 디버깅, 근본 원인 분석 및 특정 원하는 결과를 얻기 위해 반사실적 추론을 사용해야 하는 작업 계획 등 인과 추론 관련 응용 분야에 중요한 영향을 미칩니다. TempoBench를 사용하여 학습하면 오픈 소스 모델의 인과 관계 벤치마크 성능이 향상되는 동시에, 일반적인 목적, 수학 및 코드 추론 데이터 세트와 동일한 수준의 표준 벤치마크 성능을 유지할 수 있습니다. 이는 반사실적 인과 관계 추론이 기존의 추론 능력 위에 구축된 학습 가능한, 아키텍처적으로 구별되는 기능이라는 것을 시사합니다.

Original Abstract

Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern matching and forward simulation of reasoning, but underperform at counterfactual causal understanding and reasoning. We introduce TempoBench, the first formally verifiable temporal benchmark that isolates counterfactual causal attribution over execution trajectories, and we show that LLMs categorically fall back to brute-force simulation-based reasoning to solve causal reasoning problems. Built from synthesized deterministic Mealy machines, TempoBench provides an infinitely scalable corpus of trajectory-based causal reasoning problems with controllable complexity and provably correct causal labels. Frontier models reach 96% step accuracy simulating a system forward, and fall to 32% when asked which inputs were necessary for an observed output, displaying what we call the SIM/MIN gap. Our findings show that LLMs cannot reliably identify minimal necessary causes, often confusing ``possible inputs'' with ``necessary causes,'' demonstrating an inability to understand which inputs were not needed. This failure is critical for deployment in causal inference tasks such as debugging, root cause analysis, and task planning where agents must use counterfactual reasoning to plan for specific desired outcomes. We show that training on TempoBench yields a targeted gain on causal benchmarks in open-source models while matching general-purpose, math, and code reasoning datasets on standard benchmarks. This indicates counterfactual causal reasoning is a learnable, architecturally distinct capability that sits on top of existing reasoning competencies.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!