2607.01084v1 Jul 01, 2026 cs.AI

에이전트가 개방형 환경으로 일반화될 수 있는가? 도구 사용에서의 정적 학습의 취약성을 밝히다

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Zi-Jian Cheng
Zi-Jian Cheng
Citations: 17
h-index: 3
Lan-Zhe Guo
Lan-Zhe Guo
Citations: 121
h-index: 6
Weiming Wu
Weiming Wu
Citations: 2
h-index: 1
Song-Lin Lv
Song-Lin Lv
Citations: 4
h-index: 1
Ruisi Zhu
Ruisi Zhu
Citations: 16
h-index: 1

대규모 언어 모델(LLM) 에이전트는 정적인 벤치마크에서 능숙한 모습을 보이지만, 실제 시나리오에 적용할 때 사용자 질문, 도구 세트 및 상호 작용 방식의 역동성 때문에 어려움을 겪습니다. 이러한 일반화 격차를 해결하기 위해, 우리는 개방형 환경(Open-World)에서의 도구 사용 에이전트를 의미하는 'OpenAgent'라는 문제 설정을 정의합니다. 이 설정은 질문, 행동, 관찰 및 도메인 측면에서 분포 변화가 발생하는 특징을 가집니다. 이러한 영향력을 체계적으로 진단하기 위해, 우리는 세밀한 환경 변화를 네 가지 계층(인지, 상호 작용, 추론, 내재화)으로 정의한 제어된 테스트 환경을 구축하고 광범위한 실험을 수행했습니다. 분석 결과, 지도 학습(SFT) 및 강화 학습 모두를 통해 훈련된 에이전트가 개방형 환경 변화에 직면했을 때 다양한 수준의 성능 저하를 경험한다는 중요한 사실을 밝혀냈습니다. 이러한 통찰력을 바탕으로, 우리는 SFT를 위한 교란 기반 개입 전략인 'Perturbation-Augmented Fine-Tuning'을 제안합니다. 이는 에이전트의 견고성과 실제 환경에서의 유용성을 향상시키는 기반을 제공합니다. 관련 코드는 다음 주소에서 공개될 예정입니다: https://github. com/LAMDA-NeSy/OpenAgent.

Original Abstract

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions. To systematically diagnose its impact, we construct a controlled sandbox environment where we define fine-grained environmental shifts across a four-tier hierarchy, Perception, Interaction, Reasoning, and Internalization, and conduct a comprehensive series of experiments. Our analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning(SFT) and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts. Building on these insights, we propose Perturbation-Augmented Fine-Tuning, a disturbance-based intervention strategy for SFT that lays the foundation for enhancing agent robustness and utility in realistic environments. Our code will be released at: https://github. com/LAMDA-NeSy/OpenAgent.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!