2605.29927v1 May 28, 2026 cs.CL

계획 방식이 중요할까요? LLM 웹 에이전트의 계획 표현에 대한 실증적 연구

Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

Xingyu Lu
Xingyu Lu
Citations: 140
h-index: 3
Alejandra Zambrano
Alejandra Zambrano
Citations: 115
h-index: 2
Sara Vera Marjanovic
Sara Vera Marjanovic
Citations: 4
h-index: 1
Imene Kerboua
Imene Kerboua
Citations: 183
h-index: 5
Leila Kosseim
Leila Kosseim
Citations: 10
h-index: 3

최근의 발전에도 불구하고, LLM 기반 웹 에이전트는 여전히 제한적인 탐색 능력, 중요한 단계 누락, 그리고 작업 제약 조건에 대한 민감성 등의 문제점을 가지고 있습니다. 기존 연구는 이러한 실패 원인이 계획 수립 과정의 미흡함에서 비롯된다고 지적하지만, 대체 자연어 계획 표현 방식이 성능에 미치는 영향은 아직까지 제대로 조사되지 않았습니다. 본 연구에서는 PlanAhead라는 정적 계획-실행 프레임워크를 통해 계획 표현 방식이 에이전트 성능에 미치는 영향을 평가합니다. 먼저, WebArena 작업을 자동으로 3가지 난이도 수준으로 분류하여 인간의 주석 없이 일관된 난이도 등급을 부여합니다. 그런 다음, 순차적 하위 목표, 서술형, 유사 코드, 그리고 체크리스트라는 4가지 다양한 계획 표현 방식을 사용하여 난이도가 높은 작업에 대해 평가를 진행하며, OpenAI, Alibaba, Google에서 제공하는 다양한 멀티모달 LLM 기반 에이전트를 사용합니다. 무작위 변동성을 고려하기 위해, 달성률(Achievement Rate, AR)과 해결된 작업의 일관성(Solved-Task Consistency, STC)이라는 새로운 평가 지표를 도입했습니다. 연구 결과는 계획 수립 방식과 계획을 생성하는 LLM 모두가 웹 에이전트의 안정성과 작업 성공에 상당한 영향을 미친다는 것을 보여줍니다.

Original Abstract

Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints. Prior work suggests that many of these failures stem from weaknesses in planning, yet the impact of alternative natural language plan representation remains unexplored. To address this, we introduce PlanAhead, a static planner-executor framework that evaluates the impact of plan representation in agent performance. We first automatically categorize WebArena tasks into 3 difficulty levels, enabling consistent difficulty grading without human annotation. Then we systematically evaluate 4 different plan representations on the tasks categorized as hard: sequential subgoals, narrative, pseudocode, and checklist; across different families of multimodal LLM powered agents (OpenAI, Alibaba, and Google). To account for stochastic variability, we introduce two novel evaluation metrics: Achievement Rate (AR) and Solved-Task Consistency (STC). Our results show that both, the plan formulation and the underlying LLM generating the plan, significantly influence web-agent robustness and task success.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!