구체화된 에이전트 아키텍처 설계 자동화
Automating the Design of Embodied AgentArchitectures
구체화된 에이전트는 일반적으로 인식, 기억, 계획 및 행동 모듈의 수동으로 설계된 조합으로 구축됩니다. 이러한 모듈성은 방대한 아키텍처 설계 공간을 제공하지만, 현재 시스템은 여전히 연구자의 직관에 의존하여 정보가 저장되는 위치, 관찰 데이터 처리 방식, 모델 호출 연결 방식을 결정합니다. 에이전트 아키텍처 탐색(AAS)은 텍스트 도메인 에이전트에 대한 이러한 설계를 자동화하지만, 시뮬레이션 실행을 통해 인식 기반 구체화된 에이전트에 대해 체계적으로 평가되지 않았습니다. 본 연구에서는 이 기술의 적용 가능성을 조사합니다. 우리는 AgentCanvas라는 타입-그래프 런타임을 소개하며, 이는 구체화된 실행기를 편집 가능한 노드 및 와이어 프로그램으로 호스팅하고 시뮬레이터 인식을 갖춘 실행과 에피소드 수준 로그를 제공합니다. 또한 KDLoop라는 코딩 에이전트 탐색 절차를 도입하여 제안, 비평, 실험 및 증류 단계를 거치며, 정체 발생 시 반성을 유도합니다. 우리는 4가지 구체화된 실행기(시각-언어 내비게이션, 구체화된 질의 응답, 언어 기반 조작)에 대한 세 가지 AAS 변형을 평가했습니다. 결과적으로 얻은 3x4 행렬은 아키텍처 수준의 탐색이 구체화된 작업에서 배포 가능한 성공률 향상을 가져올 수 있음을 보여주지만, 한 가지 높은 점수를 받은 후보는 정보 유출 문제를 가지고 있어 채택되지 않았습니다. 동시에, 본 연구에서는 텍스트 도메인 AAS에서 감춰지는 제약 조건을 드러냅니다. 최적화 신호가 시뮬레이션 노이즈에 의해 가려질 수 있으며, 탐색이 국소적인 편집 영역에 갇힐 수 있고, 상세한 로그가 제공되더라도 에피소드 수준의 기여도 할당이 부분적으로만 나타날 수 있습니다. 이러한 결과는 구체화된 에이전트에 대한 자동 아키텍처 탐색의 잠재력과 현재 한계를 동시에 보여줍니다.
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems still rely on researcher intuition to choose where information is stored, how observations are processed, and how model calls are connected. Agent Architecture Search (AAS) automates such design for text-domain agents, but has not been systematically evaluated on perceptual embodied agents through simulator rollouts. We study this transfer. We introduce AgentCanvas, a typed-graph runtime that hosts embodied executors as editable node-and-wire programs with simulator-aware execution and episode-level logs, and KDLoop, a coding-agent search procedure that cycles through proposal, critique, experiment, and distillation, with triggered reflection after stalls. We evaluate three AAS variants across four embodied executors spanning vision-language navigation, embodied question answering, and language-conditioned manipulation. The resulting 3x4 matrix shows that architecture-level search can produce deployable and directional success-rate gains on embodied tasks, while one apparent high-scoring candidate is rejected as leak-bearing. At the same time, the experiments expose constraints that are muted in text-domain AAS: optimization signals can be masked by rollout noise, search can become trapped in local edit basins, and episode-level credit assignment only partially emerges even when detailed logs are available. These results characterize both the promise and the current limits of automated architecture search for embodied agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.