VIA: 로봇 제어를 위한 시각 인터페이스 에이전트
VIA: Visual Interface Agent for Robot Control
로봇 조작은 시각적 이해, 물리적 추론, 계획 및 폐루프 제어를 필요로 하는 복잡한 작업입니다. 범용 기반 모델(FM)은 특히 시각 및 추론 능력에서 놀라운 발전을 이루었습니다. 이러한 능력을 활용하여 일반적인 로봇 정책을 개발하기 위한 현재 방법은 일반적으로 기존 FM을 비전-언어-액션(VLA) 모델로 변환하고, 로봇 데이터를 사용하여 저수준 액션을 출력하도록 미세 조정하는 방식을 사용합니다. 그러나 VLA는 제한된 데이터 및 컴퓨팅 자원으로 인해 미세 조정을 수행해야 하므로, 종종 최첨단 FM보다 훨씬 작은 규모를 가집니다. 이는 결과적으로 VLA의 일반적인 능력을 제한합니다. 시각 인터페이스를 통해 소프트웨어를 작동하는 FM의 능력 증가에 영감을 받아, 동일한 역량이 로봇 제어에도 적용될 수 있는지 질문했습니다. 본 논문에서는 VIA(Visual Interface Agent for robot control)라는 프레임워크를 제시하며, 이는 로봇 제어를 에이전트 기반 작업으로 재구성합니다. 즉, 상용 FM 기반 에이전트가 웹 브라우저 기반 3D 인터페이스를 통해 스크린샷을 찍고, 직관적인 명령을 내리고, 결과를 관찰하고, 조작기를 제어합니다. 이 에이전트는 로봇에 특화된 미세 조정이나 특별한 상태 정보 접근 권한 없이 작동하며, 시각적 입력을 받아들이고 제한된 범위의 일반적인 도구를 사용하여 행동합니다. VIA는 에이전트의 일반적인 추론 능력, 폐루프 오류 복구 기능 및 관찰을 통해 계획하고 재계획하는 능력을 상속받습니다. VIA는 Claude Code와 Codex를 사용하여 다양한 테이블탑 조작 작업을 제로샷으로 해결했습니다. 가장 강력한 모델(Fable 5)을 사용했을 때, 세 가지 LIBERO-Goal 작업에서 96.7%의 성공률과 장기적인 레인보우 조립 작업에서 100%의 성공률을 달성했습니다. 성능은 기본 모델의 규모와 능력에 따라 향상됩니다. 이러한 결과는 최첨단 에이전트가 적절한 인터페이스를 통해 로봇 제어에 직접 적용될 수 있는 능력을 이미 보유하고 있으며, 코딩 또는 컴퓨터 사용 에이전트는 본질적으로 로봇 제어 에이전트라는 것을 시사합니다.
Robot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control. General-purpose foundation models (FMs) have grown remarkably capable of some of these, especially vision and reasoning. To leverage this for generalist robot policies, current methods typically involve converting existing FMs into vision-language-action (VLA) models by fine-tuning on robot data to output low-level actions. However, VLAs are often orders of magnitude smaller than frontier FMs given the limited data and compute available for fine-tuning, which in turn limits their general capability. Inspired by the growing ability of FMs to operate software through visual interfaces, we ask whether that same competence suffices to control a robot. We present VIA (Visual Interface Agent for robot control), a framework that recasts robot control as an agentic task: an off-the-shelf FM-powered agent drives a manipulator through a browser-based 3D interface by taking screenshots, issuing intuitive commands, observing the outcome, and adjusting. The agent receives no robot-specific fine-tuning and no access to privileged state information: it perceives visual input and acts through a small set of general tools. VIA inherits the agent's general reasoning, closed-loop error recovery, and ability to plan and re-plan from what it observes. It solves a diverse suite of tabletop manipulation tasks zero-shot with both Claude Code and Codex. With the strongest model (Fable 5) it achieves 96.7% success on three LIBERO-Goal tasks and 100% on a long-horizon rainbow assembly task. Performance improves with the scale and strength of the underlying model. These results suggest that frontier agents already possess skills that transfer directly to robot control given the right interface: your coding or computer-use agent is, in a sense, secretly a robot-control agent.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.