2607.26148v1 Jul 28, 2026 cs.RO

구체화된 에이전트가 주도하다: 최소 인터페이스 기반의 제로샷 에이전트가 시각-언어 네비게이션에서 산업 규모 정책과 경쟁한다

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

Zerui Li
Zerui Li
Citations: 113
h-index: 4
G. Zhou
G. Zhou
Citations: 860
h-index: 8
Qi Wu
Qi Wu
Citations: 293
h-index: 5
Jian Zhou
Jian Zhou
Citations: 7
h-index: 2
Sihao Lin
Sihao Lin
Citations: 471
h-index: 8
Xunyi Zhao
Xunyi Zhao
Citations: 17
h-index: 2
Jiajun Liu
Jiajun Liu
Citations: 76
h-index: 6

자율적인 구체화된 에이전트는 인식, 행동, 검증 및 여러 단계에 걸쳐 자체 수정하는 긴 의사 결정 루프를 유지해야 합니다. 현재 시스템은 작업별 워크플로우 또는 구체화된 정책을 통해 이 루프를 유지합니다. 본 연구에서는 일반적인 에이전트가 자체적으로 루프를 제어하는 '에이전틱 임베디드 제어'라는 세 번째 방식을 탐구합니다. 제로샷 네비게이션을 엄격하게 통제된 테스트 환경으로 활용하여, 단일 모노 RGB 카메라와 이산적인 행동만을 사용하여 작동하는 세 가지 소프트웨어 엔지니어링 에이전트 프레임워크를 평가했습니다. 이러한 극도로 제한된 조건에서, 기본 설정의 복제는 70.7 ± 3.5%의 성공률을 달성했으며, 'fable-5'는 최대 노력 시 78%의 성공률을 보였습니다. 학습된 웨이포인트 도구를 기본적인 기능과 함께 선택적으로 제공할 경우, 하이브리드 'fable-5' 에이전트는 기본 노력으로 76.7 ± 0.6%의 성공률을 달성했으며, 최대 노력 기반 방식보다 절반의 환경 단계를 사용하고 전체 실행 시간은 4분의 1 이하로 줄었습니다. 통제된 실험 결과, 성능 향상은 주로 모델 선택에 의해 결정되며, 프레임워크 효과는 설명적인 수준이고, 강제 웨이포인트 인터페이스는 약한 모델에는 도움이 되지만 강력한 모델의 성능을 저해할 수 있습니다. 그러나 전체적으로 더 긴 시간 범위를 갖는 작업에서는 성능이 급격히 떨어지며, 지연 시간과 컨텍스트 증가가 지속적인 작동을 제한합니다. 이러한 결과는 에이전틱 제어가 이미 제로샷 네비게이션에서 경쟁력을 갖추고 있으며, 모델, 프레임워크 및 인터페이스가 자율적인 구체화된 에이전트를 개발하기 위한 상호 보완적인 경로를 제공한다는 것을 보여줍니다.

Original Abstract

Autonomous embodied agents must sustain a long decision-making loop that involves perceiving, acting, verifying, and self-correcting over many steps. Current systems sustain this loop through task-specific workflows or embodied policies. We study a third form, agentic embodied control, in which a general-purpose agent holds the loop itself. Using zero-shot navigation as a controlled testbed, we evaluate three software-engineering agent harnesses given only a monocular RGB camera and discrete actions. Under this strictly minimal condition, replicated default-effort configurations reach 70.7$\pm$3.5% success (opus-5, mean over three runs), and fable-5 reaches 78% at maximum effort. When a trained waypoint tool is exposed alongside primitives as an optional capability, the hybrid fable-5 agent reaches 76.7$\pm$0.6% at default effort, using half the environment steps and less than one quarter of the wall time of the maximum-effort primitive run. Controlled interventions show that capability is primarily model-centered: model choice strongly changes success, harness effects are descriptive, and a forced waypoint interface helps weaker models but can hinder stronger ones. Performance nevertheless falls sharply on longer-horizon tasks, while latency and context growth limit sustained operation. These results show that agentic control is already competitive in zero-shot navigation and that models, harnesses, and interfaces offer complementary paths toward autonomous embodied agents.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!