자기 수정 웹 에이전트를 위한 검증 가능한 약속 기반 계획
Falsifiable Commitment Planning for Self-Correcting Web Agents
장기적인 목표를 가진 웹 에이전트는 종종 최종 실패 이전에 경로에서 벗어나는 경향을 보입니다. 즉, 현재 상태, 재사용 가능한 기술 또는 계획의 가정에 더 이상 사용자의 지시를 뒷받침하지 못하더라도 특정 경로는 여전히 부분적으로 타당해 보일 수 있습니다. 기존 에이전트는 계획을 세우거나, 반성을 하거나, 경험을 재사용할 수 있지만, 실행 중인 단계가 언제까지 신뢰할 수 있는지에 대한 증거는 거의 명시적으로 나타나지 않습니다. 본 논문에서는 강력한 장기 목표를 가진 웹 에이전트를 위한 검증 가능한 약속 기반 계획 프레임워크인 FCPAgent를 제안합니다. FCPAgent는 각 계획 단계를 '검증 가능한 약속 단위'(FCU)로 표현하며, 이는 재사용 가능한 기술에 근거한 하위 목표와 함께 확인 증거, 반증 증거 및 신뢰도 점수를 포함합니다. 실행은 계획-테스트-수정 루프로 구성됩니다. 하이브리드 약속 테스트 모듈은 브라우저를 수정하기 전에 후보 액션을 확인하고, 실행 후에는 관찰 결과를 확인하며, 효율성을 위해 경량 증거 매칭과 LLM 기반 진단 검증을 결합합니다. 증거가 약속을 반증하는 경우, 범위 인식 수리(scope-aware repair)는 모순을 실행, 기술 또는 계획 수준으로 제한하고 가장 적절한 부분을 수정합니다. WebArena 환경에서 FCPAgent는 가장 강력한 기준 모델에 비해 평균 성공률이 13.8% 향상되었으며, 특히 장기 목표를 가진 작업에서 큰 개선 효과를 보였습니다.
Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reuse experience, but their plans rarely specify the evidence under which an active step should still be trusted. We propose FCPAgent, a falsifiable commitment planning framework for robust long-horizon web agents. FCPAgent represents each plan step as a Falsifiable Commitment Unit (FCU): a subgoal grounded in a reusable skill, together with confirming evidence, falsifying evidence, and a confidence score. Execution is organized as a plan-test-repair loop. The hybrid commitment testing module checks candidate actions before they modify the browser and checks observations after execution; for efficiency, it combines lightweight evidence matching with LLM-based diagnostic verification. When evidence falsifies a commitment, scope-aware repair localizes the contradiction to the execution, skill, or planning level and revises the smallest adequate part. On WebArena, FCPAgent achieves a 13.8% relative improvement in average success over the strongest baseline, with especially large gains on long-horizon tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.