InternAgentHarness: LLM 에이전트 능력을 향상시키는 확장 가능한 합성 환경
InternAgentHarness: A Scalable Synthetic Environment for Enhancing LLM Agentic Abilities
대규모 언어 모델(LLM)은 점점 더 복잡한 실제 문제를 해결할 수 있는 범용 에이전트로 사용될 것으로 기대됩니다. 이러한 에이전트를 훈련하려면 상태가 변경되고 도구를 활용하는 작업과의 반복적인 상호 작용을 지원하며 검증 가능한 피드백을 제공하는 안정적이고 다양한 환경이 필요합니다. 최근의 발전에도 불구하고, 견고한 LLM 에이전트 개발은 여전히 현실적이고 확장 가능하며 실행 가능한 훈련 환경 부족으로 인해 제한됩니다. 본 논문에서는 LLM 에이전트의 능력을 향상시키기 위한 확장 가능한 합성 환경인 InternAgentHarness를 소개합니다. InternBootcamp 훈련 프레임워크~\[citep{internbootcampv1}\]를 기반으로 구축된 InternAgentHarness는 프롬프트 생성, 도구 실행, 상호 작용 제어 및 보상 계산을 통합하는 4계층 인터페이스를 통해 실행 가능한 에이전트 환경을 구현합니다. 또한, 다양한 에이전트 작업을 Bootcamp에서 훈련할 수 있는 방식으로 자동 변환하는 에이전트 하니스인 \bootcampcli를 소개합니다. InternAgentHarness는 정적인 벤치마크와 달리 평가 결과를 실질적인 개선으로 이어집니다. 관찰된 실패 사례는 체계적으로 새로운 합성 작업, 필터링된 트레이저리, 강화 학습 시뮬레이션 및 동일한 실행 가능한 인터페이스에서의 후속 재평가로 변환될 수 있습니다. InternAgentHarness를 텍스트 기반 및 비전 기반 에이전트 시나리오를 모두 포괄하는 10개의 작업 세트에 적용했습니다. Qwen3-VL-30B-A3B-Thinking 모델을 초기 모델로 사용하여 InternAgentHarness에서 수행한 지도 학습(SFT)과 강화 학습(RL)은 튜닝되지 않은 모델보다 성능이 크게 향상되었습니다. 이러한 결과는 InternAgentHarness가 확장 가능한 합성 에이전트 환경을 위한 실용적인 기반을 제공하며, 평가, 합성, 훈련 및 개선의 반복적인 주기를 통해 LLM 에이전트를 지속적으로 발전시킬 수 있음을 시사합니다.
Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated interaction with stateful, tool-augmented tasks and provide verifiable feedback. Despite recent progress, the development of robust LLM agents remains limited by the lack of realistic, scalable, and executable training environments. We present InternAgentHarness, a scalable synthetic environment for improving the agentic capabilities of LLMs. Built upon the InternBootcamp training framework~\citep{internbootcampv1}, InternAgentHarness instantiates executable agent environments through a four-layer interface that unifies prompt generation, tool execution, interaction control, and reward computation. We further introduce \bootcampcli, an agent harness that automatically converts diverse agentic tasks into a Bootcamp-trainable paradigm. Unlike static benchmarks, InternAgentHarnessmakes evaluation actionable: observed failures can be systematically converted into new synthetic tasks, filtered trajectories, reinforcement learning rollouts, and subsequent re-evaluation under the same executable interface. We instantiate InternAgentHarness on a suite of 10 tasks covering both text-only and vision-based agent scenarios. Starting from Qwen3-VL-30B-A3B-Thinking, both supervised fine-tuning (SFT) and reinforcement learning (RL) on InternAgentHarness substantially improve performance over untuned model. These results suggest that InternAgentHarness provides a practical foundation for scalable synthetic agent environments and enables the continuous improvement of LLM agents through an iterative cycle of evaluation, synthesis, training, and refinement.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.