2607.11197v1 Jul 13, 2026 cs.AI

LLM 계획에 대해 이야기할 때 우리는 무엇을 말하는가: 두 가지 구별되는 계획 능력에 대한 증거

What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities

Zhixi Cai
Zhixi Cai
Citations: 87
h-index: 6
Fucai Ke
Fucai Ke
Citations: 157
h-index: 6
Sukai Huang
Sukai Huang
Monash University
Citations: 58
h-index: 4
Gholamreza Haffari
Gholamreza Haffari
Citations: 10
h-index: 1
Chenyuan Zhang
Chenyuan Zhang
Citations: 8
h-index: 2
Naim Rastgoo
Naim Rastgoo
Citations: 0
h-index: 0
Hamid Rezatofighi
Hamid Rezatofighi
Citations: 24
h-index: 2

LLM이 다양한 계획 작업에서 불균등한 성능을 보일 때, 이러한 차이는 종종 작업의 난이도로 설명됩니다. 그러나 우리는 이 설명이 완전하지 않다고 주장합니다. 왜냐하면 작업 수준의 변동성은 단일 능력 스펙트럼에서의 차이라기보다는 뚜렷한 잠재적 계획 역량의 차이를 반영할 수 있기 때문입니다. 본 연구에서는 ACPBench-Hard 데이터셋을 사용하여 다양한 LLM 패밀리를 테스트 시간 추론 예산 변화에 따라 평가하고, 다차원 문항 반응 이론 모델을 적용하여 LLM 계획의 기반이 되는 잠재적 역량 구조를 분석합니다. 분석 결과, 계획 성능을 결정하는 두 가지 주요 차원이 밝혀졌습니다. 첫째는 운영 추론(operational reasoning)으로, 이는 로컬 액션의 적용 가능성과 즉각적인 상태 변화를 평가하는 능력입니다. 둘째는 구조적 열거(structural enumeration)로, 이는 목표 달성 가능성과 랜드마크 구조에 대해 추론하는 능력입니다. 운영 추론은 모델 크기 확장 및 더 긴 추론 과정을 통해 개선되는 반면, 구조적 열거는 상대적으로 민감하지 않습니다. 본 연구의 결과는 LLM 계획 능력을 역량 수준에서 평가하는 것을 촉구하며, 모델이 전반적으로 얼마나 개선되었는지에 대한 관심에서 벗어나 어떤 계획 역량이 어떻게, 언제, 왜 개선되는지에 초점을 맞추도록 합니다.

Original Abstract

When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete, as task-level variation may reflect distinct latent planning competencies rather than differences along a single ability spectrum. We study this question on ACPBench-Hard by evaluating multiple LLM families under varying test-time reasoning budgets and applying a multidimensional item response theory model to uncover the latent competency structure underlying LLM planning. The analysis reveals two principal dimensions that shape planning performance: operational reasoning, the ability to evaluate local action applicability and immediate state transitions, and structural enumeration, the ability to reason about goal reachability and landmark structure. Operational reasoning improving under model scaling and longer reasoning traces, while structural enumeration remains comparatively insensitive. Our findings motivate competency-level evaluation of LLM planning, shifting the focus from whether models improve overall to which planning competencies improve, under what conditions, and why.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!