Dynamo: 시각-언어 에이전트를 위한 동적 스킬-도구 진화
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
시각 추론 능력을 향상시키기 위해 일반적으로 시각-언어 모델(VLM)을 재학습하거나 수동으로 설계된 프롬프트 및 도구를 사용합니다. 본 연구에서는 가중치 업데이트 없이 고정된 VLM을 적응시키는 훈련 불필요한 프레임워크인 Dynamo를 제시합니다. 작은 크기의 레이블이 지정된 학습 데이터 세트를 사용하여, 에이전트는 자신의 정답 및 오답 시도를 검토하고 인지적 병목 현상을 해결하기 위한 재사용 가능한 추론 스킬과 인식적 문제를 해결하기 위한 실행 가능한 시각 도구라는 두 가지 상호 보완적인 기능을 발전시킵니다. 생성된 각 도구는 해당 도구를 언제 호출해야 하는지를 명시하는 스킬과 함께 제공되며, 이러한 모든 기능 유형은 지속적인 라이브러리에 축적됩니다. 네 가지 시각 추론 벤치마크와 다섯 가지 VLM 아키텍처에서 Dynamo는 모든 20개의 모델-벤치마크 설정에서 직접 추론 성능을 향상시켰습니다(평균 +5.6% 정확도). 도구 세트가 미리 제공되는 경우, 프레임워크는 각 도구를 언제 호출해야 하는지를 학습하며, 단계별 도구 선택은 모든 테스트된 아키텍처에서 성능 향상을 가져왔습니다. 특정 작업에 최적화된 강화 학습(VTool-R1, DeepEyes)과 비교했을 때, Dynamo는 훨씬 적은 연산량으로 65~99%의 강화 학습 성능 격차를 줄이며, 필요한 경우 강화 학습과 함께 사용될 수 있습니다.
Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects its own correct and incorrect attempts and evolves two complementary capabilities: reusable reasoning skills for cognitive bottlenecks, and executable visual tools for perceptual ones. Each generated tool is paired with a skill that specifies when to invoke it, and both capability types accumulate in a persistent library. Across four visual reasoning benchmarks and five VLM backbones, Dynamo improves direct inference on all 20 model--benchmark settings (avg. +5.6 acc). When the tool set is given in advance, the framework learns when to call each tool, and per-step tool choice improves on every tested backbone. Against task-specific RL (VTool-R1, DeepEyes), Dynamo closes 65--99% of the RL gap at a fraction of the compute, and combines additively with RL when available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.