2606.30185v1 Jun 29, 2026 cs.AI

Dynamo: 시각-언어 에이전트를 위한 동적 스킬-도구 진화

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Haoxuan Ma
Haoxuan Ma
Citations: 35
h-index: 4
Mengyu Zhou
Mengyu Zhou
Citations: 66
h-index: 3
Xiaoxi Jiang
Xiaoxi Jiang
Citations: 142
h-index: 6
Guanjun Jiang
Guanjun Jiang
Citations: 117
h-index: 5
Yutao Sun
Yutao Sun
Citations: 98
h-index: 5
Mingshuai Chen
Mingshuai Chen
Citations: 42
h-index: 4
Dexin Wang
Dexin Wang
Citations: 46
h-index: 2
Li Xu
Li Xu
Citations: 646
h-index: 4
Yanting Miao
Yanting Miao
Citations: 13
h-index: 2
Tiancheng Zhao
Tiancheng Zhao
Citations: 381
h-index: 7
Lei Lv
Lei Lv
Citations: 0
h-index: 0

시각 추론 능력을 향상시키기 위해 일반적으로 시각-언어 모델(VLM)을 재학습하거나 수동으로 설계된 프롬프트 및 도구를 사용합니다. 본 연구에서는 가중치 업데이트 없이 고정된 VLM을 적응시키는 훈련 불필요한 프레임워크인 Dynamo를 제시합니다. 작은 크기의 레이블이 지정된 학습 데이터 세트를 사용하여, 에이전트는 자신의 정답 및 오답 시도를 검토하고 인지적 병목 현상을 해결하기 위한 재사용 가능한 추론 스킬과 인식적 문제를 해결하기 위한 실행 가능한 시각 도구라는 두 가지 상호 보완적인 기능을 발전시킵니다. 생성된 각 도구는 해당 도구를 언제 호출해야 하는지를 명시하는 스킬과 함께 제공되며, 이러한 모든 기능 유형은 지속적인 라이브러리에 축적됩니다. 네 가지 시각 추론 벤치마크와 다섯 가지 VLM 아키텍처에서 Dynamo는 모든 20개의 모델-벤치마크 설정에서 직접 추론 성능을 향상시켰습니다(평균 +5.6% 정확도). 도구 세트가 미리 제공되는 경우, 프레임워크는 각 도구를 언제 호출해야 하는지를 학습하며, 단계별 도구 선택은 모든 테스트된 아키텍처에서 성능 향상을 가져왔습니다. 특정 작업에 최적화된 강화 학습(VTool-R1, DeepEyes)과 비교했을 때, Dynamo는 훨씬 적은 연산량으로 65~99%의 강화 학습 성능 격차를 줄이며, 필요한 경우 강화 학습과 함께 사용될 수 있습니다.

Original Abstract

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects its own correct and incorrect attempts and evolves two complementary capabilities: reusable reasoning skills for cognitive bottlenecks, and executable visual tools for perceptual ones. Each generated tool is paired with a skill that specifies when to invoke it, and both capability types accumulate in a persistent library. Across four visual reasoning benchmarks and five VLM backbones, Dynamo improves direct inference on all 20 model--benchmark settings (avg. +5.6 acc). When the tool set is given in advance, the framework learns when to call each tool, and per-step tool choice improves on every tested backbone. Against task-specific RL (VTool-R1, DeepEyes), Dynamo closes 65--99% of the RL gap at a fraction of the compute, and combines additively with RL when available.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!