시간 퍼즐을 활용한 반복적인 시간 추론 능력 측정
Measuring Iterative Temporal Reasoning with Time Puzzles
본 연구에서는 반복적인 시간 추론 능력을 평가하기 위한 제약 기반의 날짜 추론 과제인 '시간 퍼즐(Time Puzzles)'을 소개합니다. 각 퍼즐은 사실 기반의 시간 정보와 (다양한 문화권의) 달력 정보를 결합하며, 하나 또는 여러 개의 유효한 해답 날짜를 가질 수 있습니다. 또한, 제어되고 동적이며 지속적인 평가를 위해 알고리즘적으로 생성됩니다. 13개의 다양한 LLM(Large Language Models)을 대상으로 실험한 결과, '시간 퍼즐'은 각 모델의 반복적인 시간 추론 능력을 효과적으로 구분했으며, 도구 없이 풀기에는 여전히 어려운 과제임이 드러났습니다. GPT-5는 49.3%의 정확도를 기록했지만, 다른 모든 모델은 31% 미만의 성능을 보였습니다. 웹 검색을 활용하면 상당한 성능 향상을 얻을 수 있으며, 코드 인터프리터를 사용하면 결과가 혼합되었지만, 모든 모델은 명시적인 날짜로 제약 조건을 다시 작성했을 때 훨씬 더 나은 성능을 보였습니다. 이는 도구를 활용하는 능력에 대한 격차를 보여줍니다. 전반적으로 '시간 퍼즐'은 도구 지원 반복적인 시간 추론 능력을 진단하는 간단하고 비용 효율적인 방법입니다.
We introduce Time Puzzles, a constraint-based date inference task for evaluating iterative temporal reasoning. Each puzzle combines factual temporal anchors with (cross-cultural) calendar relations, admits one or multiple valid solution dates, and is algorithmically generated for controlled, dynamic, and continual evaluation. Across 13 diverse LLMs, Time Puzzles well distinguishes their iterative temporal reasoning capabilities and remains challenging without tools: GPT-5 reaches only 49.3% accuracy and all other models stay below 31%, despite the dataset's simplicity. Web search consistently yields substantial gains and using code interpreter shows mixed effects, but all models perform much better when constraints are rewritten with explicit dates, revealing a gap in reliable tool use. Overall, Time Puzzles presents a simple, cost-effective diagnostic for tool-augmented iterative temporal reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.